When Does GraphRAG Help? A Controlled Study of Evidence Localization in Long Annual Reports

Research
RAG
Information Retrieval
Evaluation
When Does GraphRAG Help? A Controlled Study of Evidence Localization in Long Annual Reports

85.9%

Object Recall@10

91.8%

Page Recall@10

48,292

Graph nodes

Tech Stack

Python
GraphRAG
RAG
Information Retrieval
BM25
E5
Cross-encoder
Reranking
Pandas
Bootstrap
Statistics

Description

This project explores graph-based evidence retrieval from long annual reports. The graph connects related evidence through shared entities and financial metrics. As first author and experimental lead, I designed the GraphRAG approach and led the experiments and evaluation.

  • Built a GraphRAG system to find supporting evidence in long annual reports using document structure, entities, and financial metrics.
  • Compared standard retrieval with graph-based retrieval, finding that graph expansion improved evidence-object recovery while graph paths added page coverage.
  • Analyzed which graph relations helped or harmed retrieval through edge ablations, paired bootstrap tests, and a held-out-year robustness check.

Research Questions

RQ1

Which structural relations are useful for evidence retrieval, and which introduce noise?

RQ2

What do selected-graph expansion and graph-path retrieval contribute beyond a strong hybrid retriever under controlled conditions?

RQ3

How do structure-aware retrieval methods behave on multi-hop and visual/layout-sensitive questions?

Method overview

GraphRAG as Controlled Evidence Navigation

Standard retrieval finds evidence mainly through text similarity. Our GraphRAG approach also uses document structure, connecting evidence through shared pages, entities, and financial metrics. The goal is to test whether these graph connections can recover useful evidence that standard retrieval misses.

The system combines evidence from three sources: hybrid retrieval provides the original results, selected-graph expansion finds related evidence from those results, and graph paths independently find evidence using cues from the question. All candidates are then combined and reranked.

Simplified controlled GraphRAG fusion pipeline
Simplified web illustration of the controlled fusion pipeline.

RQ1. Which Graph Connections Help?

The graph connects evidence in four ways: evidence can appear on the same page, mention the same entity, refer to the same financial metric, or appear on adjacent pages.

We test these connections one at a time to see which ones help GraphRAG find the right evidence and which ones introduce noise.

Simplified typed metadata graph schema
Evidence is connected through document structure, shared entities, and shared financial metrics.
Edge-type ablation for Object Recall at 10
Each relation is removed in turn to measure how it affects retrieval.

The results show that same-entity connections are the most useful, while same-metric connections provide a smaller benefit. Same-page connections have limited impact. Adjacent-page connections are harmful because they bring in distracting evidence from nearby pages.

Key finding: More graph connections are not always better. What matters is connecting evidence that is meaningfully related.

Does This Finding Hold in Later Reports?

Our main analysis shows that adjacent-page links can introduce noise and hurt retrieval. We repeat the comparison on reports from 2022–2024 to see whether the same finding holds in later years.

MetricWith adjacent-page linksWithout adjacent-page links
Object Recall@1070.0%84.7%
MRR0.6330.729

Removing adjacent-page links again improves retrieval, suggesting that this effect is consistent across different years within the Equinor reports. This check does not test generalization to other companies or document collections.

RQ2. What Does Each Graph Source Add?

We compare the original hybrid retrieval with two graph-based methods: Selected Graph and Graph Paths. The original hybrid results are kept, and each graph method adds additional evidence before the final ranking.

MethodObject Recall@10Page Recall@10
Hybrid E583.8%90.8%
+ Selected Graph85.6%90.3%
+ Graph Paths85.0%91.4%
+ Graph + Paths85.9%91.8%

The two graph methods help in different ways. Selected Graph improves the retrieval of specific evidence, increasing Object Recall@10 from 83.8% to 85.6%. Graph Paths mainly help find relevant pages: when added to Selected Graph, Page Recall@10 increases from 90.3% to 91.8%.

Together, Graph + Paths gives the best overall coverage, reaching 85.9% Object Recall@10 and 91.8% Page Recall@10.

RQ3. Where Does Retrieval Still Fail?

Even with graph-based retrieval, two challenges remain: finding all the evidence needed for multi-hop questions and finding the exact evidence inside visually complex pages.

Multi-Hop Questions

Some questions require evidence from multiple places in the reports. Finding only one of them is not enough.

Retrieval ResultHybrid E5+ Graph + Paths
At least one evidence object found100.0%100.0%
All evidence objects found67.8%72.2%
All evidence pages found71.1%74.4%

Both methods can always find some relevant evidence. The harder problem is finding all the evidence needed for the question. Adding Graph and Paths improves this from 67.8% to 72.2%.

Key finding: Multi-hop retrieval still struggles to find all required evidence, even when some relevant evidence is easy to find.

Visual and Layout Questions

For questions involving tables, figures, or page layout, the system may find the correct page but still miss the specific evidence on that page.

Retrieval ResultHybrid E5+ Graph + Paths
Evidence object found41.1%50.0%
Evidence page found73.3%78.9%

This gap is still large. With Graph and Paths, the system reaches the correct page in 78.9% of cases, but finds the specific evidence in only 50.0%.

Key finding: For visual and layout questions, finding the right page is much easier than finding the exact evidence inside it.