FinRAG-Equinor: A Reliability-Audited Benchmark for Evidence Localization in Longitudinal Annual Reports

Research
RAG
Information Retrieval
Evaluation
Data Quality
Human Evaluation
FinRAG-Equinor: A Reliability-Audited Benchmark for Evidence Localization in Longitudinal Annual Reports

Tech Stack

Python
RAG
Information Retrieval
BM25
E5
Cross-encoder
Reranking
Pandas
Benchmarking
Bootstrap
Cohen's Kappa

Description

FinRAG-Equinor is a reliability-audited benchmark for evidence-grounded RAG on long annual reports. It is built from 15 Equinor/Statoil annual reports (2010–2024) and includes 720 QA items with traceable report-, page-, and object-level evidence.

The benchmark evaluates whether RAG systems can identify the right report, locate the relevant page, and find the exact evidence needed to answer each question.

As the first author and experimental lead, I led the design and implementation of the benchmark and retrieval evaluation pipeline, from evidence construction and QA development to systematic evaluation of sparse, dense, hybrid, hierarchical, and reranked retrieval methods.

  • Built a traceable evidence corpus from long annual-report PDFs, preserving report-, page-, and object-level provenance.
  • Constructed and audited the benchmark across factual, numerical, table-grounded, temporal, multi-hop, visual/layout, policy, causal, and unanswerable cases.
  • Evaluated retrieval and end-to-end QA across sparse, dense, hybrid, hierarchical, and reranked approaches, examining how evidence localization affects answer quality.
  • Analyzed retrieval failures at report, page, object, ranking, and multi-hop levels, and connected retrieval quality to end-to-end answer accuracy.

Research Questions

RQ1.Does finding the right report and the right page make it easier to find the exact evidence?

RQ2.Where does retrieval fail?

RQ3.Does retrieving better evidence lead to more accurate answers?

From Annual-Report PDFs to Traceable Evidence

We turn annual-report PDFs into retrieval units that remain linked to their original report, page, and location, then use them to build an audited QA benchmark.

01

Annual reports

15 reports · 4,369 pages

02

Layout objects

100,150 extracted objects

03

Retrieval corpus

41,736 paragraphs, headings, and tables

04

Benchmark

720 QA items · 660 answerable

Benchmark Composition and Reliability

The benchmark covers nine question types. We manually checked all 720 QA items and separately audited PDF extraction and QA quality.

Question typeItems
Direct factual90
Numerical extraction90
Definition / policy60
Causal explanation75
Temporal, year-specific75
Table-grounded90
Multi-hop90
Visual / chart layout90
Unanswerable60

95 / 100

Audited PDF pages usable for RAG

100%

Agreement on overall QA usability

99%

Agreement on answer correctness

98%

Agreement on evidence support

RQ1. Does finding the right report and the right page make it easier to find the exact evidence?

RQ1 tests two things: whether searching only the correct report helps, and whether page-sized context makes the exact evidence easier to retrieve.

EXPERIMENT 1

Does searching the correct report help?

BM25 settingObject R@10Page R@10
Unrestricted51.5%60.3%
Reference-year filtered75.6%85.0%

Object Recall@10 requires the exact reference evidence object to appear in the top 10. Page Recall@10 requires a result from the same report and page. Reference-year filtering uses the benchmark's known report year, so it is a controlled oracle setting rather than a learned report router.

Reference-year filtering raises exact-object Recall@10 from 51.5% to 75.6%, showing that cross-year competition is a major source of retrieval error.

EXPERIMENT 2

Does retrieving more context help?

Using the correct report year, we compare objects, pages, page windows, and object windows.

Retrieval unitChunk Recall@10Page Recall@10
Object BM25-year75.6%85.0%
Page BM25-year90.5%90.5%
Page-window87.6%87.6%
Object-window88.3%91.7%

An object is one paragraph, heading, or table. A page contains all retrieval objects on one PDF page. A page-window combines up to three consecutive pages: the previous, current, and next pages. An object-window combines eight neighboring objects within the same report, with a two-object overlap.

Chunk Recall@10 measures whether at least one of the top 10 retrieved chunks contains an annotated reference evidence object. Page Recall@10 measures whether at least one of the top 10 chunks covers the same report and page as a reference evidence object.

Whole pages achieve the highest Chunk Recall@10 at 90.5%. Larger windows do not improve chunk retrieval, although object windows achieve the highest Page Recall@10 at 91.7%.

Finding the correct report produces the largest improvement in exact evidence retrieval. Larger retrieval units improve evidence coverage, but finding a page that contains the evidence is not the same as retrieving the exact supporting object.

RQ2. Where does retrieval fail?

Finding the correct report is only the first step. Even within the correct report, retrieval can still fail in different ways.

Does it find the right evidence?

We give each retrieval method the correct report and examine its top-10 results. This shows whether it finds the exact evidence or retrieves evidence from the right page or report but misses the correct object.

BM25 represents sparse keyword retrieval, BGE-M3 represents dense semantic retrieval, and BM25 + E5 represents hybrid retrieval.

Retrieval methodExact evidence foundRight page, wrong objectRight report, wrong page
BM2575.6%9.4%9.7%
BGE-M382.1%6.8%8.2%
BM25 + E583.8%7.0%5.9%

Dense and hybrid retrieval find the exact evidence more often than BM25. BM25 + E5 achieves the highest exact-evidence rate and the lowest wrong-page rate, while BGE-M3 has the lowest same-page wrong-object rate.

Note: Rates use all 660 answerable questions. Adjacent-page cases are not shown.

Is the right evidence ranked first?

Even when the correct evidence is retrieved, it may not appear at the top. We therefore compare how often the exact evidence is ranked first before and after reranking.

MeasureBefore rerankingAfter reranking
Exact evidence ranked first40.2%63.6%

Reranking moves the correct evidence higher, but cannot recover evidence that was not retrieved.

Can it find all the evidence?

Some questions require multiple pieces of evidence. For these multi-hop questions, finding one relevant piece is not enough; the retrieval method needs to find all the evidence required to answer the question.

Multi-hop outcomeRecall@10
At least one required evidence item91.1%
All required evidence53.3%

Finding some evidence is much easier than finding all the evidence needed.

Note: Multi-hop evaluation covers 90 questions.

Retrieval can fail at three stages: finding the right evidence, ranking it high enough, and retrieving the complete evidence set.

RQ3. Does retrieving better evidence lead to more accurate answers?

To answer RQ3, we use an end-to-end experiment: the same answer generator is given different evidence, and we compare the resulting answer accuracy.

Answer accuracy on 660 questions

Evidence given to the generatorAnswer accuracy
Question only (closed-book)2.9%
BM25-year evidence58.5%
Hybrid + reranked evidence70.9%
Annotated reference evidence82.0%

Question only (closed-book) means that the model receives no retrieved evidence. BM25-year and Hybrid + reranked use automatically retrieved evidence. Annotated reference evidence gives the model the evidence linked to each question in the benchmark.

Hybrid + reranked evidence improves answer accuracy by 12.4 percentage points over BM25-year, but remains 11.1 points below annotated reference evidence. Better evidence leads to more accurate answers; the remaining gap shows that both retrieval and generation can still improve.

RQ2's multi-hop results explain part of this gap: retrieval may find relevant evidence while still missing information required for a complete answer.

Scope

This is a controlled longitudinal study of one company and 15 English annual reports. Cross-company, multilingual, and fully multimodal evaluation remain future work.