LongFinRAG: A Benchmark for Finding the Right Evidence in Annual Reports

Tech Stack
Description
LongFinRAG is a reliability-audited benchmark for evidence-grounded RAG on long annual reports. It is built from 15 Equinor/Statoil annual reports (2010–2024) and includes 720 QA items with traceable report-, page-, and object-level evidence.
The benchmark evaluates whether RAG systems can identify the right report, locate the relevant page, and find the exact evidence needed to answer each question.
As the first author and experimental lead, I led the design and implementation of the benchmark and retrieval evaluation pipeline, from evidence construction and QA development to systematic evaluation of sparse, dense, hybrid, hierarchical, and reranked retrieval methods.
- Built a traceable evidence corpus from long annual-report PDFs, preserving report-, page-, and object-level provenance.
- Constructed and audited the benchmark across factual, numerical, table-grounded, temporal, multi-hop, visual/layout, policy, causal, and unanswerable cases.
- Evaluated retrieval and end-to-end QA across sparse, dense, hybrid, hierarchical, and reranked approaches, examining how evidence localization affects answer quality.
- Analyzed retrieval failures at report, page, object, ranking, and multi-hop levels, and connected retrieval quality to end-to-end answer accuracy.
Research Questions
RQ1.How do the year filter and retrieval-unit size affect evidence retrieval?
RQ2.Where and how does retrieval fail?
RQ3.Does better evidence lead to better answers?
From Annual-Report PDFs to Traceable Evidence
We turn annual-report PDFs into retrieval units that remain linked to their original report, page, and location, then use them to build an audited QA benchmark.
01
Annual reports
15 reports · 4,369 pages
02
Layout objects
100,150 extracted objects
03
Retrieval corpus
41,736 paragraphs, headings, and tables
04
Benchmark
720 QA items · 660 answerable
Benchmark Composition and Reliability
The benchmark covers nine question types. We manually checked all 720 QA items and separately audited PDF extraction and QA quality.
| Question type | Items |
|---|---|
| Direct factual | 90 |
| Numerical extraction | 90 |
| Definition / policy | 60 |
| Causal explanation | 75 |
| Temporal, year-specific | 75 |
| Table-grounded | 90 |
| Multi-evidence | 90 |
| Layout-sensitive | 90 |
| Unanswerable | 60 |
95 / 100
Audited PDF pages usable for RAG
100%
Agreement on overall QA usability
99%
Agreement on answer correctness
98%
Agreement on evidence support
How We Evaluate LongFinRAG
We evaluate LongFinRAG in two ways: evidence retrieval and end-to-end question answering. The first tests whether the system can find the right evidence, while the second tests whether better evidence leads to better answers.
RQ1. How do the year filter and retrieval-unit size affect evidence retrieval?
We test two things. First, we compare Without year filter with With reference-year filter. Second, we compare objects, pages, page-windows, and object-windows as retrieval units.
EXPERIMENT 1
Does the reference-year filter help?
Without year filter searches all 15 reports. With reference-year filter searches only the correct report.
Find the exact evidence
Object Recall@10
BM25
BGE-M3
BM25 + E5
Find the correct page
Page Recall@10
BM25
BGE-M3
BM25 + E5
Without year filter → With reference-year filter
Exact evidence: the paragraph or table that contains the answer.
Correct page: the page where that evidence appears.
The reference-year filter improves both results for all three methods. BM25 + E5 gives the highest results: 84.5% for exact evidence and 91.4% for the correct page.Takeaway: The reference-year filter makes it easier to find both the correct page and the exact evidence.
EXPERIMENT 2
Does more context help?
We use the With reference-year filter setting and compare four retrieval units.
| Retrieval unit | Evidence found @10 | Correct page found @10 |
|---|---|---|
| Object | 76.5% | 85.8% |
| Page | 91.2% | 91.2% |
| Page-window | 88.5% | 88.5% |
| Object-window | 89.2% | 92.4% |
What does each unit contain?
- Object: one paragraph, heading, or table
- Page: all objects on one PDF page
- Page-window: up to three nearby pages
- Object-window: eight nearby objects
Metrics
- Evidence found: a top-10 result contains the needed evidence
- Correct page found: a top-10 result reaches the correct page
Whole-page retrieval works best for finding the evidence (91.2%). Object-window retrieval works best for finding the correct page (92.4%).Takeaway: One page works best for finding the exact evidence, while adding nearby objects helps find the correct page. Adding more pages does not improve the results.
RQ2. Where and how does retrieval fail?
Retrieval can fail in three ways: it may search in the wrong place, rank the correct evidence too low, or miss some of the required evidence.
EXPERIMENT 1
Where does retrieval fail?
We compare the top-10 results with the reference evidence.
| Retrieval outcome | Without year filter | With reference-year filter |
|---|---|---|
| Exact object | 52.4% | 76.5% |
| Right page, wrong object | 8.8% | 9.2% |
| Nearby page | 6.7% | 5.0% |
| Right report, wrong page | 23.6% | 9.2% |
| Wrong report | 8.5% | 0% |
If a question has multiple reference evidence objects, finding at least one counts as an exact-object hit. The five categories do not overlap, so each row adds up to 100%.
Takeaway: Providing the year information helps BM25 find the right report, but it does not always find the right page or object.
EXPERIMENT 2
Can reranking fix ranking errors?
The correct evidence may be in the top 10 but ranked too low. Two cross-encoders, MiniLM and BGE, reorder the same top-10 results.
Object Recall@1 · Higher is better
Object Recall@1 measures how often the first result is an exact reference object. MRR also improves with both rerankers.
Reranking improves the top result for both candidate sets. MiniLM is slightly better in both cases.Takeaway: Reranking cannot find evidence that is not already in the top 10.
EXPERIMENT 3
Can retrieval find all the evidence?
Some questions need more than one piece of evidence.
We look at two things:
- Any evidence: at least one required evidence item is found
- All evidence: all required evidence items are found
Any vs. All Evidence Recall@10
Results on 90 multi-evidence questions. All methods use reference-year filtering; hybrid results use RRF fusion.
Takeaway: Finding at least one piece of evidence is common, but finding all the evidence is much harder.
RQ3. Does better evidence lead to better answers?
We use the same questions and the same answer model, but provide different evidence.
Answer accuracy on 660 questions
| Evidence given to the model | Accuracy |
|---|---|
| Question only | 3.2% |
| BM25-year evidence | 58.9% |
| Hybrid + reranked evidence | 71.4% |
| Annotated reference evidence | 82.4% |
Question only: No retrieved evidence
BM25-year / Hybrid + reranked: Automatically retrieved evidence
Annotated reference evidence: Evidence linked to each question
All four settings correctly abstain on all 60 unanswerable questions.
Takeaway
Better evidence leads to better answers.
Even with annotated reference evidence, some questions remain difficult, especially layout-sensitive questions.
What We Learned
LongFinRAG shows that finding the right report is not enough. A reliable RAG system also needs to:
- •Find the right page and exact evidence.
- •Find all the evidence when a question needs more than one piece.
- •Use the evidence to produce the correct answer.
Scope: LongFinRAG studies one company and 15 English annual reports. Cross-company, multilingual, and multimodal evaluation remain future work.