I built this to learn how to tell whether an AI agent is actually doing its job, not just producing answers that sound right, using flight planning for two real trips as the test case.
Langfuse is an observability tool: it collects what each run did and turns it into latency, token usage, cost and scores you can compare. It never watches the agent, the agent reports to it, so a step nobody wrote a line for is a step nobody can see.
You can't open the model. You can see everything crossing the line, which is enough to measure a change, compare two models on the same work, and keep the version that behaves better.
A trace is the record of a single run: every model turn, every tool call, what went in, what came back and how long each took. It exists because the agent's own code opens it and fills it in. Here is that code, with the parts that matter numbered.
1with tracer.trace("plan-trip", session_id=run_id, tags=[scenario, model]):23 4with tracer.observation(as_type="generation", name="decide-next-step") as gen: response = client.beta.messages.create(model=model, tools=TOOL_DEFS, ...)5 gen.update(output=response.content, usage_details={"input": n_in, "output": n_out})6
with block is the record. Opening it starts the trace and the clock. When the
indented code finishes, the block closes and the duration is stored. Start and stop are just the edges of a
block of code you chose.plan-trip, which is what makes a hundred of them comparable.generation tells the interface this step was a model call, so it shows the prompt, the reply and
the tokens.Six stages, each watched by a different check. That is what makes a wrong step visible even when the final answer reads well.
Two real trips on live Google Flights data, plus stress tests on synthetic fares and injected API failures. Every model saw identical data, so differences are the model, not the fares.
The two trips above plus thirteen stress tests.
Swipe the table sideways to see all models. Tap a result for the full run.
Averages across every scenario. Hover a bar for exact numbers.
The traces showed a fix worth making, and the same runs measured whether it worked. That loop is the reason to build the eval.
Same loop as the change above, one step earlier: the traces raised a question I hadn't thought to ask, and the same eval can settle which answer is worth shipping. Proposed, not built.
The API keeps no memory between turns, so each turn re-sends the whole conversation: the system prompt, the tool definitions, and every search result collected so far. The reply the user reads is the small part.
Nothing in the final answers hinted at this. It only showed up because every turn's token counts were recorded.
How I'd settle it: run the same 15 scenarios before and after, and compare cost, latency and all six checks. A change that saves money and breaks grounding is not a saving.
One thing first. The runs currently record input and output tokens only. Caching adds separate counts for cache writes and cache reads, and plain input tokens exclude both, so switching it on without recording them would make the bill look better than it is. Fix the accounting, then measure.