VerdictBench: Benchmarking Long Context and Retrieval-Augmented Generation on Indonesian Constitutional Court Verdicts
An empirical benchmark comparing Long Context, Dense RAG, and Multi-Stage RAG on 50 Indonesian Constitutional Court verdicts and 300 human-reviewed QA pairs. Phase 2 shows Long Context and Dense RAG are statistically tied on gold-evidence faithfulness across Gemini 2.5 Flash and GPT-4o Mini, while Dense RAG is 16-25x cheaper and avoids the long-verdict non-response failures seen with Long Context. Multi-Stage RAG underperforms both baseline methods under oracle gold-evidence evaluation, making the study a cautionary result for adding query rewriting, hybrid search, metadata filters, and reranking without domain validation.
Proof points
- LC ~= Dense RAG
- 24.8x Cheaper
- 300 QA Pairs
Technologies
Python, Gemini 2.5 Flash, GPT-4o Mini, FAISS, RAG, BM25, IndoBERT, NLP
Links