LongFinRAG: A Benchmark for Finding the Right Evidence in Annual Reports

Research
RAG
Information Retrieval
Evaluation
Data Quality
Human Evaluation
LongFinRAG: A Benchmark for Finding the Right Evidence in Annual Reports

Tech Stack

Python
RAG
Information Retrieval
BM25
E5
Cross-encoder
Reranking
Pandas
Benchmarking
Bootstrap
Cohen's Kappa

Description

LongFinRAG is a reliability-audited benchmark for evidence-grounded RAG on long annual reports. It is built from 15 Equinor/Statoil annual reports (2010–2024) and includes 720 QA items with traceable report-, page-, and object-level evidence.

The benchmark evaluates whether RAG systems can identify the right report, locate the relevant page, and find the exact evidence needed to answer each question.

As the first author and experimental lead, I led the design and implementation of the benchmark and retrieval evaluation pipeline, from evidence construction and QA development to systematic evaluation of sparse, dense, hybrid, hierarchical, and reranked retrieval methods.

  • Built a traceable evidence corpus from long annual-report PDFs, preserving report-, page-, and object-level provenance.
  • Constructed and audited the benchmark across factual, numerical, table-grounded, temporal, multi-hop, visual/layout, policy, causal, and unanswerable cases.
  • Evaluated retrieval and end-to-end QA across sparse, dense, hybrid, hierarchical, and reranked approaches, examining how evidence localization affects answer quality.
  • Analyzed retrieval failures at report, page, object, ranking, and multi-hop levels, and connected retrieval quality to end-to-end answer accuracy.

Research Questions

RQ1.How do the year filter and retrieval-unit size affect evidence retrieval?

RQ2.Where and how does retrieval fail?

RQ3.Does better evidence lead to better answers?

From Annual-Report PDFs to Traceable Evidence

We turn annual-report PDFs into retrieval units that remain linked to their original report, page, and location, then use them to build an audited QA benchmark.

01

Annual reports

15 reports · 4,369 pages

02

Layout objects

100,150 extracted objects

03

Retrieval corpus

41,736 paragraphs, headings, and tables

04

Benchmark

720 QA items · 660 answerable

Benchmark Composition and Reliability

The benchmark covers nine question types. We manually checked all 720 QA items and separately audited PDF extraction and QA quality.

Question typeItems
Direct factual90
Numerical extraction90
Definition / policy60
Causal explanation75
Temporal, year-specific75
Table-grounded90
Multi-evidence90
Layout-sensitive90
Unanswerable60

95 / 100

Audited PDF pages usable for RAG

100%

Agreement on overall QA usability

99%

Agreement on answer correctness

98%

Agreement on evidence support

How We Evaluate LongFinRAG

We evaluate LongFinRAG in two ways: evidence retrieval and end-to-end question answering. The first tests whether the system can find the right evidence, while the second tests whether better evidence leads to better answers.

LongFinRAG evaluation pipeline showing evidence retrieval for RQ1 and RQ2 and end-to-end question answering for RQ3

RQ1. How do the year filter and retrieval-unit size affect evidence retrieval?

We test two things. First, we compare Without year filter with With reference-year filter. Second, we compare objects, pages, page-windows, and object-windows as retrieval units.

EXPERIMENT 1

Does the reference-year filter help?

Without year filter searches all 15 reports. With reference-year filter searches only the correct report.

Find the exact evidence

Object Recall@10

52.4
76.5

BM25

71.5
83.0

BGE-M3

67.1
84.5

BM25 + E5

Find the correct page

Page Recall@10

61.2
85.8

BM25

79.1
89.8

BGE-M3

75.5
91.4

BM25 + E5

Without year filterWith reference-year filter

Without year filter → With reference-year filter

Exact evidence: the paragraph or table that contains the answer.
Correct page: the page where that evidence appears.

The reference-year filter improves both results for all three methods. BM25 + E5 gives the highest results: 84.5% for exact evidence and 91.4% for the correct page.Takeaway: The reference-year filter makes it easier to find both the correct page and the exact evidence.

EXPERIMENT 2

Does more context help?

We use the With reference-year filter setting and compare four retrieval units.

Retrieval unitEvidence found @10Correct page found @10
Object76.5%85.8%
Page91.2%91.2%
Page-window88.5%88.5%
Object-window89.2%92.4%

What does each unit contain?

  • Object: one paragraph, heading, or table
  • Page: all objects on one PDF page
  • Page-window: up to three nearby pages
  • Object-window: eight nearby objects

Metrics

  • Evidence found: a top-10 result contains the needed evidence
  • Correct page found: a top-10 result reaches the correct page

Whole-page retrieval works best for finding the evidence (91.2%). Object-window retrieval works best for finding the correct page (92.4%).Takeaway: One page works best for finding the exact evidence, while adding nearby objects helps find the correct page. Adding more pages does not improve the results.

RQ2. Where and how does retrieval fail?

Retrieval can fail in three ways: it may search in the wrong place, rank the correct evidence too low, or miss some of the required evidence.

EXPERIMENT 1

Where does retrieval fail?

We compare the top-10 results with the reference evidence.

Retrieval outcomeWithout year filterWith reference-year filter
Exact object52.4%76.5%
Right page, wrong object8.8%9.2%
Nearby page6.7%5.0%
Right report, wrong page23.6%9.2%
Wrong report8.5%0%

If a question has multiple reference evidence objects, finding at least one counts as an exact-object hit. The five categories do not overlap, so each row adds up to 100%.

Takeaway: Providing the year information helps BM25 find the right report, but it does not always find the right page or object.

EXPERIMENT 2

Can reranking fix ranking errors?

The correct evidence may be in the top 10 but ranked too low. Two cross-encoders, MiniLM and BGE, reorder the same top-10 results.

Retrieve top 10→MiniLM / BGE rerank→Check rank 1
BeforeMiniLMBGE
0
20
40
60
80
41.1
64.7
63.6
49.8
68.3
68.0
BM25BM25 + BGE

Object Recall@1 · Higher is better

Object Recall@1 measures how often the first result is an exact reference object. MRR also improves with both rerankers.

Reranking improves the top result for both candidate sets. MiniLM is slightly better in both cases.Takeaway: Reranking cannot find evidence that is not already in the top 10.

EXPERIMENT 3

Can retrieval find all the evidence?

Some questions need more than one piece of evidence.

We look at two things:

  • Any evidence: at least one required evidence item is found
  • All evidence: all required evidence items are found

Any vs. All Evidence Recall@10

Any evidenceAll evidence
0
25
50
75
100
91.1
53.3
95.6
63.3
95.6
61.1
98.9
74.4
100.0
67.8
Object BM25Page BM25Object Qwen3BM25 + BGE hybridBM25 + E5 hybrid

Results on 90 multi-evidence questions. All methods use reference-year filtering; hybrid results use RRF fusion.

Takeaway: Finding at least one piece of evidence is common, but finding all the evidence is much harder.

RQ3. Does better evidence lead to better answers?

We use the same questions and the same answer model, but provide different evidence.

Answer accuracy on 660 questions

Evidence given to the modelAccuracy
Question only3.2%
BM25-year evidence58.9%
Hybrid + reranked evidence71.4%
Annotated reference evidence82.4%

Question only: No retrieved evidence

BM25-year / Hybrid + reranked: Automatically retrieved evidence

Annotated reference evidence: Evidence linked to each question

All four settings correctly abstain on all 60 unanswerable questions.

Takeaway

Better evidence leads to better answers.

Even with annotated reference evidence, some questions remain difficult, especially layout-sensitive questions.

What We Learned

LongFinRAG shows that finding the right report is not enough. A reliable RAG system also needs to:

  • •Find the right page and exact evidence.
  • •Find all the evidence when a question needs more than one piece.
  • •Use the evidence to produce the correct answer.

Scope: LongFinRAG studies one company and 15 English annual reports. Cross-company, multilingual, and multimodal evaluation remain future work.