Multi-Agent Guardrails Evaluation FastAPI + SSE

Atlas — agentic trip planner

Six specialist agents research live flights, real hotel rates, weather, and things to do, then reconcile the plan against a budget. Every number traces back to a tool call, arithmetic is Python, and uncertainty is labelled rather than smoothed over. Built entirely on free-tier APIs — a constraint that shaped the whole design.

Role
Solo · end-to-end
Agents
6 · 4 in parallel
Tests
608 offline + 12 evals
Budget
Free tier only

01 Overview

Six specialist agents research live flights, real hotel rates, weather, and things to do, then reconcile the result against a budget — and the model is never trusted with a fact or a sum.

Most LLM travel demos ask a model what it remembers about Lisbon. This one asks tools, then checks the answer against the payload it came from. Prices are verified against the tool response that produced them. Attractions come from OpenStreetMap, and the itinerary agent cannot schedule anything outside the candidate list it was given. Cost arithmetic is Python — a budget agent that mis-adds is worse than no budget agent at all.

It runs entirely on free-tier APIs, which is not a footnote: the daily token ceiling shaped the architecture, from key rotation to cache TTLs to what gets recomputed on a retry.

02 Results

6
agents — 4 in parallel, 2 downstream, plus deterministic synthesis
608
tests, ~6s, fully offline
87.2s → 1.4s
cold run vs resumed run — successful agents are never re-paid for
12
deterministic eval scorers, 3 of them marked critical

A run takes 60–150s on free-tier models, so the browser watches the agents work over SSE rather than showing a spinner — and REQUIRED/OPTIONAL badges make it visible while running whether a failure will degrade the plan or end it.

The budget verdict is arithmetic over the agents' own figures, so it cannot be flattering. When a trip genuinely does not fit the budget, the plan says so first.

03 Architecture

Request
TripRequest
dates · travellers · budget
Fan out
flight · hotels · weather · places
4 agents in parallel
Compose
itinerary
candidates only
Price
budget
Python arithmetic
Render
synthesize
no LLM at all
Stream
SSE to browser
live agent status

synthesize is deliberately not an agent: it formats data the agents already produced, so it cannot invent a price.

Free-tier data sources
Google Flights / Hotels via SerpApiTravelpayoutsNominatim + Open-MeteoOverpass (OSM) + WikivoyageGroq models
1

Narrow tools per agent

The flight agent has find_airports and search_flights and nothing else — it has no way to look up a hotel, so it cannot wander outside its job. Tool input is validated by JSON Schema; tool output by Pydantic, because an unchecked read turns upstream API drift into silently wrong numbers.

JSON Schema inPydantic outleast-privilege tools
2

Two-phase agent loop

Phase one binds tools and lets the model call them until it stops asking. Phase two unbinds them and forces a schema against a minimal prompt plus a tool-result digest. Split because doing both at once makes the model choose between a tool call and an answer — on free-tier models that reliably produces neither.

tool phaseforced-schema phasetoken-budget aware
3

Provenance checking

Every scheduled activity is matched against the candidate list that produced it; every fare and rate against the tool payload. Findings are reported as warnings and never silently corrected — a wrong auto-fix is worse than a visible caveat. Money is parsed in Python, because letting a model convert "$148" fed straight into the budget.

price fidelitycandidate matchingwarnings not fixes
4

Honest labelling, checked

A trip 40 days out has no forecast, so the plan says "climate normals, not a forecast". Round-trip fares say so. Estimated entry fees say "estimated". Wrong-airport detection looks IATA codes up rather than recalling them — a wrong-but-real code returns valid fares for the wrong city, and nothing downstream can catch that. Closed-day detection parses OSM opening_hours narrowly and returns unknown rather than guessing.

climate normals labelledIATA validatedopening_hours parsed
5

Failure containment

Three loop limits, because they fail differently: never stops asking, never converges, re-asks the same question. Per-node timeouts with backoff that refuses to sleep past a deadline. Criticality tiers: weather failing degrades the plan and says what was lost; flight failing fails the run with a non-zero exit code, so a script cannot mistake it for success.

MAX_ROUNDS / CALLS / REPEATSper-node timeoutsREQUIRED vs OPTIONAL
6

Living inside a free tier

Groq's free tier enforces a tokens-per-day ceiling that appears in no response header — only in the 429 body. So agents map to an ordered chain of keys and rotate on a 429 rather than sleeping, concurrent agents never share a primary key, and cache TTLs follow the cost of refreshing (hotels 12h, metered; flights 4h, keyless) rather than volatility alone.

key rotation on 429single-flight lockingtiered TTLs

04 Evaluation

Evals, not just unit tests

608 unit tests check whether the code is correct. A 12-case golden set with 12 deterministic scorers checks whether the output is good — groundedness, price fidelity, honest labelling, day coverage, budget adherence, geographic coherence.

Three critical scorers

Some failures invalidate a plan rather than degrade it. Those scorers are marked critical, so a run that fails them is not reported as a partial success.

Reliability across repeats

--repeat 3 runs the golden set multiple times, because a non-deterministic system that passes once has not been shown to pass.

Cost tier in the eval

LLM calls, tokens, tool calls, retries, latency, and cache hit rate are scored alongside quality — a plan that is correct and unaffordable is not a passing result.

Checkpointed resume

Results are checkpointed per trip, so a rerun skips agents that already succeeded: 87.2s cold, 1.4s warm, identical plan. Only failures cost tokens twice.

Deterministic where possible

Trip length in nights, the free/paid distinction for places, cost arithmetic, and plan rendering are all Python. Models are used for judgement, never for facts or sums.

05 Stack

Agents
6 specialist agentstwo-phase tool loopnarrow per-agent tool sets
Serving
FastAPIServer-Sent EventsCLI entry point
Data
SerpApi (Google Flights / Hotels)TravelpayoutsOpen-MeteoNominatimOverpass (OSM)Wikivoyage
Reliability
ordered key chain + 429 rotationsingle-flight cache lockingtiered TTLsper-trip checkpoints
Quality
608 offline tests12-case golden eval setdeterministic scorersPydantic contracts

The full source is on GitHub

Every number on this page is reproducible from the repository.