GraphScout
An agentic multi-hop research assistant built as a measurement study: three retrieval arms over the same 30k-passage corpus, scored with bootstrap confidence intervals by a versioned eval harness and a calibrated LLM judge, traced per request, and gated in CI. Agentic retrieval lifts recall@8 from 0.454 to 0.838 and all-support@8 from 14% to 65% — and the more useful finding is that answer F1 barely moves, so retrieval was never the binding constraint.
01 Overview
Multi-hop questions — "which university did the founder of the company that acquired DeepMind attend?" — break naive RAG, because the second hop is invisible from the surface question. Plenty of projects assert that graph RAG or an agent fixes it. GraphScout measures how much each fix actually buys, with confidence intervals.
Three arms share one corpus, one eval harness, and one set of metrics: a dense-vector baseline, a graph-RAG arm (hybrid retrieval → rerank → knowledge graph expansion), and an agentic arm that decomposes the question and routes sub-questions across vector, graph, and community tools before verifying its answer.
The infrastructure is the point as much as the result: a versioned eval harness, a calibrated LLM-as-judge, per-request tracing in Langfuse, and a CI gate that blocks quality regressions. Total infrastructure cost: $0. Total LLM spend: ~$30, kept there by model tiering, batch processing, prompt caching, and response caching.
02 Results
All three arms on the same corpus, split and commit — dev set, n=150, k=8, MuSiQue, bootstrap 95% CIs:
| arm | recall@8 | all-support@8 | answer EM | answer F1 | $ / run |
|---|---|---|---|---|---|
| 1 · baseline — dense, single-shot | 0.454 [.406,.505] | 0.140 [.087,.200] | 0.080 | 0.149 | $0.82 |
| 2 · graph RAG — hybrid → rerank → graph-expand | 0.518 [.471,.570] | 0.220 [.160,.287] | 0.140 | 0.219 | $0.00* |
| 3 · agentic — decompose → route → verify | 0.838 [.798,.875] | 0.647 [.573,.720] | 0.127 | 0.231 | $9.28 |
*The graph arm's scored run hit a fully warm response cache. Every row traces to a committed, provenance-stamped run record — all three at the same git SHA, corpus and split.
Scored as chunk-level coverage: a gold paragraph counts as retrieved only if some chunk covering it lands in the top-8, over a 50-candidate pool.
Agentic retrieval is what closes the multi-hop gap. Recall@8 nearly doubles over the dense baseline, and all-support@8 — every gold chunk for a question present in the top-8, which is what multi-hop answering actually requires — goes from 14% to 65%. Neither single-shot arm could touch that, and it took decomposition and iteration (6.2 tool rounds per question on average) rather than a smarter single query.
And retrieval was never the binding constraint on answers. Answer F1 moves 0.149 → 0.231 while recall moves 0.39 points, and exact match is actually lower than the graph arm (0.127 vs 0.140). Grounding sits at 0.613 and only 65.3% of agent runs finish cleanly. So the honest reading is that the agent now finds the evidence and still often fails to compose it into the gold answer — the next work is answer synthesis and verification, not more retrieval. That half is stated because a 2× recall headline sitting on top of a flat F1 is exactly the result a portfolio is tempted to crop.
Within single-shot retrieval, an earlier free ablation isolates why arm 2 helps at all: both single-shot arms plateau at the same ~0.62 recall@50, so hybrid retrieval and reranking don't find evidence dense search missed — they pull known-good evidence higher (MRR 0.68 → 0.77). Single-shot retrieval gets all gold chunks into the top-8 for only 19% of questions, which is what makes the agentic arm's 65% the interesting number.
03 Architecture
Every prompt is versioned and hash-tracked outside the code; each eval arm is one YAML file, so the config diff is the experiment.
Full-article corpus, chunked
~3,900 Wikipedia articles chunked into 30k passages, indexed once with a corpus fingerprint so every later result names the corpus it was measured on.
Hybrid retrieval
Dense pgvector HNSW search and BM25 fused with reciprocal-rank fusion, then reranked by a local bge cross-encoder — no hosted reranking API, no per-query cost.
Closed-schema knowledge graph
Entity and relation extraction against a closed schema (an ADR records why open-ended extraction was rejected), alias resolution, Neo4j loading, and Leiden community summaries. Retrieval gains a graph-expansion stage that reaches neighbouring passages.
Agentic arm
The agent decomposes a multi-hop question, routes each sub-question across vector, graph, and community tools, then verifies the assembled answer. Bounded concurrency and response caching keep per-question spend measurable rather than open-ended.
Measurement infrastructure
A versioned eval harness, a calibrated LLM-as-judge, per-request cost and latency tracing in Langfuse, and a CI regression gate scored against committed baseline run records. 139 unit tests pass; 7 integration tests run with the containers up.
04 Engineering decisions
Tiered models, batched calls
Haiku / Sonnet / Opus are used where each is worth its price, with the Batches API and prompt caching on top, and every request's cost recorded — which is how a three-arm study lands at ~$30 total.
Prompts are artifacts, not strings
Every prompt lives in prompts/, versioned and hash-tracked, never inline in code — so a result can be tied to the exact prompt that produced it.
The config diff is the experiment
One YAML per arm on a shared base. Comparing arms means reading a diff, not trusting that two code paths were otherwise identical.
Retracted findings stay visible
The git history includes a reverted reranker conclusion and the defects behind it. A measurement pipeline is only worth something if it can overturn its own earlier claim.
Confidence intervals on everything
Bootstrap 95% CIs on every metric, because a 5-point recall difference on n=150 is not automatically a difference at all.
Regressions blocked in CI
A quality gate runs the harness against baseline records on every change, so a refactor that quietly degrades retrieval fails the build instead of shipping.
05 Stack
The full source is on GitHub
Every number on this page is reproducible from the repository.