Evaluation¶
Vector Graph RAG is evaluated on three standard multi-hop QA benchmarks used in the HippoRAG papers.
Datasets¶
| Dataset | Description | Hop Count | Source |
|---|---|---|---|
| MuSiQue | Multi-hop questions requiring 2–4 reasoning steps | 2–4 hops | Paper |
| HotpotQA | Wikipedia-based multi-hop QA | 2 hops | Paper |
| 2WikiMultiHopQA | Cross-document reasoning over Wikipedia | 2 hops | Paper |
Recall@5 measures how much of the ground-truth supporting evidence appears in the top five retrieved results, independently of answer generation.
Historical Results¶
These are the original three-dataset results, published before Jev was introduced into the project in September 2026.3 The latest two-stage Jev evaluation builds on this work with higher recall on MuSiQue and 2Wiki.
Recall@5 vs. Naive RAG¶
| Method | MuSiQue | HotpotQA | 2WikiMultiHopQA | Average |
|---|---|---|---|---|
| Naive RAG | 55.6% | 90.8% | 73.7% | 73.4% |
| Vector Graph RAG | 73.0% | 96.3% | 94.1% | 87.8% |
| Improvement | +31.4% | +6.1% | +27.7% | +19.6% |
Vector Graph RAG improves over Naive RAG by 19.6% in relative average Recall@5, with the largest gains on MuSiQue and 2WikiMultiHopQA.
Comparison with State-of-the-Art¶
| Method | MuSiQue | HotpotQA | 2WikiMultiHopQA | Average |
|---|---|---|---|---|
| HippoRAG (ColBERTv2)1 | 51.9% | 77.7% | 89.1% | 72.9% |
| IRCoT + HippoRAG1 | 57.6% | 83.0% | 93.9% | 78.2% |
| NV-Embed-v22 | 69.7% | 94.5% | 76.5% | 80.2% |
| HippoRAG 22 | 74.7% | 96.3% | 90.4% | 87.1% |
| Vector Graph RAG | 73.0% | 96.3% | 94.1% | 87.8% |
The historical evaluation reached 87.8% average Recall@5, leading on 2Wiki while trailing HippoRAG 2 on MuSiQue. Two-stage Jev now improves both: 76.78% on MuSiQue and 95.35% on 2Wiki, ahead of the compared baselines on each dataset. See the Jev results below for the full comparison.
Methodology¶
We reuse HippoRAG’s pre-extracted triplets to keep the graph input consistent across these experiments. Retrieval and reranking configurations are described in the corresponding results.
Evaluation Setup¶
flowchart LR
T["HippoRAG's\npre-extracted\ntriplets"] --> I["Index into\nMilvus"]
I --> Q["Run benchmark\nqueries"]
Q --> R["Check if gold\npassages in top-5"]
R --> M["Compute\nRecall@5"]
- Triplets: Use HippoRAG's pre-extracted
(subject, predicate, object)triplets from each benchmark dataset - Indexing: Build the vector knowledge graph in Milvus using these triplets
- Querying: Run all benchmark questions through the query pipeline
- Scoring: Check whether the ground-truth supporting passages appear in the top-5 retrieved results
Reproduction¶
Full reproduction steps are available in the evaluation directory:
# Clone the repository
git clone https://github.com/zilliztech/vector-graph-rag.git
cd vector-graph-rag
# See evaluation instructions
cat evaluation/README.md
See evaluation/README.md for detailed instructions.
Jev Reranker Evaluation¶
We began evaluating Jev in September 2026, following its availability, and expanded to the full two-stage evaluation below in October 2026.
Reaching the quality–latency Pareto frontier¶
Two-stage Jev achieves the highest Recall@5 among the compared methods on both datasets, averaging 86.07% across MuSiQue and 2Wiki, 1,000 questions each. Under the reference latency estimates, it reaches the quality–latency Pareto frontier with 3.13 seconds of additional model-call time.
The pipeline first selects relations, then reranks their source passages together with direct vector-search candidates. The second stage judges full passages, so the final ordering can use evidence beyond the relation text. It is a separate recipe and dataset scope from the historical three-dataset average above.
| Method | MuSiQue | 2Wiki | Average |
|---|---|---|---|
| Naive RAG · BGE-large-en-v1.5 | 58.03 | 72.73 | 65.38 |
| Naive RAG · ColBERTv2 | 49.20 | 68.20 | 58.70 |
| Naive RAG · NV-Embed-v2 | 69.70 | 76.50 | 73.10 |
| HippoRAG · ColBERTv2 | 51.90 | 89.10 | 70.50 |
| IRCoT + HippoRAG | 57.60 | 93.90 | 75.75 |
| HippoRAG 2 | 74.70 | 90.40 | 82.55 |
| Vector Graph RAG + GPT-4o-mini | 64.41 | 91.35 | 77.88 |
| Vector Graph RAG + GPT-5-mini | 73.32 | 93.80 | 83.56 |
| Vector Graph RAG + Jev · relations only | 69.30 | 90.78 | 80.04 |
| Vector Graph RAG + Jev · two-stage | 76.78 | 95.35 | 86.07 |

Compared with relation-only Jev, the second stage adds 6.03 percentage points of average Recall@5 for about 0.85 seconds more model-call time. It also improves average recall by 2.51 points over Vector Graph RAG + GPT-5-mini and 3.52 points over HippoRAG 2.
On the plotted Pareto frontier, no alternative offers both at least the same recall and no more additional model-call time, with a strict improvement in either. Two-stage Jev is the frontier's highest-quality option; direct retrieval and relation-only Jev offer lower-latency choices at lower recall.
The horizontal axis excludes embedding, retrieval and answer generation. Jev uses recorded mean request durations; generative-model points use estimated range midpoints. The dotted Pareto frontier is conditional on those estimates. Naive RAG adds no judgment call, hence zero additional time.
See the complete evaluation and reproduction instructions for the dataset checks, cached GPT replay, one historical 2Wiki fallback row, score normalization and timing sample sizes. Enable the implementation through the reranking guide.
-
HippoRAG: Neurobiologically Inspired Long-Term Memory for LLMs (NeurIPS 2024) ↩↩
-
From RAG to Memory: Non-Parametric Continual Learning for LLMs (2025) ↩↩
-
The HippoRAG authors resampled HotpotQA between HippoRAG and HippoRAG 2, so the historical HotpotQA results use different question samples. The original three-dataset results are retained for reference. This does not affect the new Jev comparison, which uses MuSiQue and 2Wiki only. ↩