Trip Agent Evals

Evaluating a flight‑search agent

I built this to learn how to tell whether an AI agent is actually doing its job, not just producing answers that sound right, using flight planning for two real trips as the test case.

Every model, side by side

How Langfuse works

Langfuse is an observability tool: it collects what each run did and turns it into latency, token usage, cost and scores you can compare. It never watches the agent, the agent reports to it, so a step nobody wrote a line for is a step nobody can see.

Visible · crosses the API boundary
  • Everything sent in: system prompt, conversation, tool definitions
  • Everything that comes back: text, tool requests, stop reason
  • Every tool result, because your own code produced it
  • Tokens, latency and cost for each turn
Not visible · inside the model
  • The computation itself
  • Why it picked this tool over another
  • The raw reasoning (summaries are a self-report, not a readout)

You can't open the model. You can see everything crossing the line, which is enough to measure a change, compare two models on the same work, and keep the version that behaves better.

What's a trace

One run, recorded by the code that runs it

A trace is the record of a single run: every model turn, every tool call, what went in, what came back and how long each took. It exists because the agent's own code opens it and fills it in. Here is that code, with the parts that matter numbered.

1with tracer.trace("plan-trip", session_id=run_id, tags=[scenario, model]):23
    4with tracer.observation(as_type="generation", name="decide-next-step") as gen:
        response = client.beta.messages.create(model=model, tools=TOOL_DEFS, ...)5
        gen.update(output=response.content, usage_details={"input": n_in, "output": n_out})6
  1. 1The with block is the record. Opening it starts the trace and the clock. When the indented code finishes, the block closes and the duration is stored. Start and stop are just the edges of a block of code you chose.
  2. 2The name. What this run is called in the interface. Every run of this agent is a plan-trip, which is what makes a hundred of them comparable.
  3. 3Session and tags. Labels for finding it later: which scenario, which model, which eval run it belonged to. This is what lets you ask for "every honeymoon run on Haiku" instead of scrolling.
  4. 4A smaller block per step. One per model turn, one per tool call. Because they open while the run's block is still open, they nest underneath it, which is why a trace reads as a tree. generation tells the interface this step was a model call, so it shows the prompt, the reply and the tokens.
  5. 5The actual work. This line is the API call, and Langfuse has nothing to do with it. Delete the tracing lines around it and the agent behaves exactly the same; you simply have no record.
  6. 6What gets recorded, including the tokens. The reply, plus the input and output token counts the API returned. These are exact, not estimated: the same numbers the bill is based on, which is where every cost on this page comes from. Leave this line out and the turn shows a duration and nothing else.

What the agent actually does

Six stages, each watched by a different check. That is what makes a wrong step visible even when the final answer reads well.

What we tested on

Two real trips on live Google Flights data, plus stress tests on synthetic fares and injected API failures. Every model saw identical data, so differences are the model, not the fares.

All 15 scenarios, every model

The two trips above plus thirteen stress tests.

Swipe the table sideways to see all models. Tap a result for the full run.

Where the time and tokens go

Averages across every scenario. Hover a bar for exact numbers.

Run time for every scenario, by model

What the eval changed

The traces showed a fix worth making, and the same runs measured whether it worked. That loop is the reason to build the eval.

Where would the next saving come from?

Same loop as the change above, one step earlier: the traces raised a question I hadn't thought to ask, and the same eval can settle which answer is worth shipping. Proposed, not built.

What the data showed
The question it raised

Why is the input so big, and which part of it is avoidable?

The API keeps no memory between turns, so each turn re-sends the whole conversation: the system prompt, the tool definitions, and every search result collected so far. The reply the user reads is the small part.

Nothing in the final answers hinted at this. It only showed up because every turn's token counts were recorded.

What I'd try, in order of expected payoff
  1. Prompt caching. Mark the stable prefix so repeat reads bill at a fraction of the normal input price. It targets the re-sent share directly, which is the largest block.
  2. Smaller tool results. The biggest single jump is the flight search payload. Returning fewer offers, or fewer fields per offer, shrinks it once and then again on every later turn that resends it.
  3. Context editing. Clear old tool results from the history once the agent has used them, so the last turn isn't still carrying the first turn's raw payload.
  4. Fewer turns. Every tool call costs a round trip: the model asks, the code runs the tool, the model is called again to see the result. One turn here does almost nothing yet pays for a full resend of everything before it. Whether two of those steps can share a turn without breaking "check before you book" is a question for the eval, not a guess.

How I'd settle it: run the same 15 scenarios before and after, and compare cost, latency and all six checks. A change that saves money and breaks grounding is not a saving.

One thing first. The runs currently record input and output tokens only. Caching adds separate counts for cache writes and cache reads, and plain input tokens exclude both, so switching it on without recording them would make the bill look better than it is. Fix the accounting, then measure.

Run diagram

Claude turnTool call