FinRAG-Equinor: A Reliability-Audited Benchmark for Evidence Localization in Longitudinal Annual Reports
Tech Stack
Description
FinRAG-Equinor is a reliability-audited benchmark for evidence-grounded RAG on long annual reports. It is built from 15 Equinor/Statoil annual reports (2010–2024) and includes 720 QA items with traceable report-, page-, and object-level evidence.
The benchmark evaluates whether RAG systems can identify the right report, locate the relevant page, and find the exact evidence needed to answer each question.
As the first author and experimental lead, I led the design and implementation of the benchmark and retrieval evaluation pipeline, from evidence construction and QA development to systematic evaluation of sparse, dense, hybrid, hierarchical, and reranked retrieval methods.
- Built a traceable evidence corpus from long annual-report PDFs, preserving report-, page-, and object-level provenance.
- Constructed and audited the benchmark across factual, numerical, table-grounded, temporal, multi-hop, visual/layout, policy, causal, and unanswerable cases.
- Evaluated retrieval and end-to-end QA across sparse, dense, hybrid, hierarchical, and reranked approaches, examining how evidence localization affects answer quality.
- Analyzed retrieval failures at report, page, object, ranking, and multi-hop levels, and connected retrieval quality to end-to-end answer accuracy.
Research Questions
RQ1.Does finding the right report and the right page make it easier to find the exact evidence?
RQ2.Where does retrieval fail?
RQ3.Does retrieving better evidence lead to more accurate answers?
From Annual-Report PDFs to Traceable Evidence
We turn annual-report PDFs into retrieval units that remain linked to their original report, page, and location, then use them to build an audited QA benchmark.
01
Annual reports
15 reports · 4,369 pages
02
Layout objects
100,150 extracted objects
03
Retrieval corpus
41,736 paragraphs, headings, and tables
04
Benchmark
720 QA items · 660 answerable
Benchmark Composition and Reliability
The benchmark covers nine question types. We manually checked all 720 QA items and separately audited PDF extraction and QA quality.
95 / 100
Audited PDF pages usable for RAG
100%
Agreement on overall QA usability
99%
Agreement on answer correctness
98%
Agreement on evidence support
RQ1. Does finding the right report and the right page make it easier to find the exact evidence?
RQ1 tests two things: whether searching only the correct report helps, and whether page-sized context makes the exact evidence easier to retrieve.
EXPERIMENT 1
Does searching the correct report help?
| BM25 setting | Object R@10 | Page R@10 |
|---|---|---|
| Unrestricted | 51.5% | 60.3% |
| Reference-year filtered | 75.6% | 85.0% |
Object Recall@10 requires the exact reference evidence object to appear in the top 10. Page Recall@10 requires a result from the same report and page. Reference-year filtering uses the benchmark's known report year, so it is a controlled oracle setting rather than a learned report router.
Reference-year filtering raises exact-object Recall@10 from 51.5% to 75.6%, showing that cross-year competition is a major source of retrieval error.
EXPERIMENT 2
Does retrieving more context help?
Using the correct report year, we compare objects, pages, page windows, and object windows.
| Retrieval unit | Chunk Recall@10 | Page Recall@10 |
|---|---|---|
| Object BM25-year | 75.6% | 85.0% |
| Page BM25-year | 90.5% | 90.5% |
| Page-window | 87.6% | 87.6% |
| Object-window | 88.3% | 91.7% |
An object is one paragraph, heading, or table. A page contains all retrieval objects on one PDF page. A page-window combines up to three consecutive pages: the previous, current, and next pages. An object-window combines eight neighboring objects within the same report, with a two-object overlap.
Chunk Recall@10 measures whether at least one of the top 10 retrieved chunks contains an annotated reference evidence object. Page Recall@10 measures whether at least one of the top 10 chunks covers the same report and page as a reference evidence object.
Whole pages achieve the highest Chunk Recall@10 at 90.5%. Larger windows do not improve chunk retrieval, although object windows achieve the highest Page Recall@10 at 91.7%.
RQ2. Where does retrieval fail?
Finding the correct report is only the first step. Even within the correct report, retrieval can still fail in different ways.
Does it find the right evidence?
We give each retrieval method the correct report and examine its top-10 results. This shows whether it finds the exact evidence or retrieves evidence from the right page or report but misses the correct object.
BM25 represents sparse keyword retrieval, BGE-M3 represents dense semantic retrieval, and BM25 + E5 represents hybrid retrieval.
| Retrieval method | Exact evidence found | Right page, wrong object | Right report, wrong page |
|---|---|---|---|
| BM25 | 75.6% | 9.4% | 9.7% |
| BGE-M3 | 82.1% | 6.8% | 8.2% |
| BM25 + E5 | 83.8% | 7.0% | 5.9% |
Dense and hybrid retrieval find the exact evidence more often than BM25. BM25 + E5 achieves the highest exact-evidence rate and the lowest wrong-page rate, while BGE-M3 has the lowest same-page wrong-object rate.
Note: Rates use all 660 answerable questions. Adjacent-page cases are not shown.
Is the right evidence ranked first?
Even when the correct evidence is retrieved, it may not appear at the top. We therefore compare how often the exact evidence is ranked first before and after reranking.
| Measure | Before reranking | After reranking |
|---|---|---|
| Exact evidence ranked first | 40.2% | 63.6% |
Reranking moves the correct evidence higher, but cannot recover evidence that was not retrieved.
Can it find all the evidence?
Some questions require multiple pieces of evidence. For these multi-hop questions, finding one relevant piece is not enough; the retrieval method needs to find all the evidence required to answer the question.
| Multi-hop outcome | Recall@10 |
|---|---|
| At least one required evidence item | 91.1% |
| All required evidence | 53.3% |
Finding some evidence is much easier than finding all the evidence needed.
Note: Multi-hop evaluation covers 90 questions.
RQ3. Does retrieving better evidence lead to more accurate answers?
To answer RQ3, we use an end-to-end experiment: the same answer generator is given different evidence, and we compare the resulting answer accuracy.
Answer accuracy on 660 questions
| Evidence given to the generator | Answer accuracy |
|---|---|
| Question only (closed-book) | 2.9% |
| BM25-year evidence | 58.5% |
| Hybrid + reranked evidence | 70.9% |
| Annotated reference evidence | 82.0% |
Question only (closed-book) means that the model receives no retrieved evidence. BM25-year and Hybrid + reranked use automatically retrieved evidence. Annotated reference evidence gives the model the evidence linked to each question in the benchmark.
RQ2's multi-hop results explain part of this gap: retrieval may find relevant evidence while still missing information required for a complete answer.
Scope
This is a controlled longitudinal study of one company and 15 English annual reports. Cross-company, multilingual, and fully multimodal evaluation remain future work.