How to Write Evals for AI Agents: A Step-by-Step Guide

September 15, 2026
How to Write Evals for AI Agents

Most teams that ship an AI agent test it the way they'd test a chatbot: run some prompts, eyeball the answers, ship it. That works until the agent starts calling tools, taking multiple steps, and failing in ways a single right-or-wrong answer can't catch — looping on a step, calling the right tool with the wrong arguments, or technically completing the task in a way that breaks something else. Writing real evals is how you catch that before your users do. Here's how to actually build them.

Step 1: Know the Difference Between a Benchmark and an Eval

A benchmark is a fixed, public task set with a leaderboard — it tells you how a model compares to other models on a standardized problem. An eval is something you build to test your own agent against your own task, on your own data. You cite a benchmark; you build a framework. If your only testing plan is “check how it does on a public benchmark,” you're comparing models, not verifying your agent actually works.

This distinction matters more than it sounds. In July 2026, OpenAI audited SWE-Bench Pro — a widely cited coding benchmark — and found that roughly 30% of its 731 tasks were broken: overly strict tests that rejected correct solutions, underspecified prompts, and tasks that passed on incomplete fixes. OpenAI retracted its own earlier recommendation to use it. If a widely respected benchmark can be a third broken, your own evals — built around your actual task — matter more than borrowing someone else's leaderboard.

Step 2: Decide Which Layer You're Actually Testing

Agent evaluation has three layers, and most teams only build the first one:

•         Final-answer evaluation — score the last message against an expected result. Necessary, but the answer can be right while the path to it was slow, expensive, or unsafe.

•         Trajectory evaluation — score the sequence of steps that produced the answer: did it call the right tools with the right arguments, and did it loop or recover cleanly from a bad tool call?

•         Online evaluation — score live production traffic in real time, using a fast classifier rather than a slow, expensive LLM-as-judge call on every single turn.

Start with final-answer evals because they're the easiest to build, but budget for trajectory evals before you ship anything that takes more than one or two steps — that's where most real agent failures actually live.

Step 3: Build a Golden Dataset From Real Traces, Not Just Synthetic Cases

A golden dataset is the set of representative tasks you'll score your agent against every time you change a prompt, a tool, or a model. Synthetic test cases are a fine starting point, but the highest-value entries come from real production traces — actual conversations where your agent looped, called the wrong tool, or frustrated a user. Pull those into your dataset as regression tests so a fixed bug can't silently come back.

If you're building this alongside a coding agent specifically, our AI Coding Agents Explained guide and Best AI Agent Framework for Python cover the frameworks these datasets typically plug into.

Step 4: Pick a Scorer — and Know Where It Breaks

Deterministic checks (file exists, regex match, exact string comparison) are fast and cheap but only work for narrowly defined tasks. LLM-as-judge scoring handles fuzzier tasks but is slow, expensive at scale, and vulnerable to reward hacking — where an agent (or a model being trained against your eval) learns to satisfy the scorer's letter rather than its intent.

Cursor's own writeup on its Bugbot code-review tool is a useful real-world example: instead of scoring whether the agent's suggestion looked plausible, they built their primary metric around post-merge signal — using a separate AI check at merge time to confirm a flagged bug was actually fixed by the developer, validated with human spot-checks. That single change, plus 40 subsequent experiments, moved their resolution rate from 52% to over 70%. The lesson generalizes: score outcomes that are hard to game, not surface plausibility.

Step 5: Wire Evals Into CI/CD So Regressions Get Caught Automatically

An eval you run manually before a big release is better than nothing, but an eval that runs on every pull request is what actually prevents regressions. Most of the tools below support this directly — either as a native CI integration or, in DeepEval's case, as a pytest assertion that runs alongside your existing test suite.

Tool Best For Pricing Standout Detail
LangSmith Teams already on LangChain/LangGraph Free (5k traces/mo); Plus $39/seat/mo SOC 2 Type II at its $39 tier — the lowest compliance entry point in the category
Braintrust Product teams shipping fast, dataset-first workflows Free Starter; Pro $249/mo flat-rate Sandboxed custom Python scorers no other platform currently offers
Arize Phoenix Open-source production tracing + eval in one stack Free, self-hosted (Elastic License) OpenTelemetry-native; parent company Arize was announced for acquisition by Dynatrace in August 2026
Promptfoo Security and red-team testing of agents Free, open-source 500+ built-in attack vectors; acquired by OpenAI in March 2026
DeepEval CI-integrated testing without a hosted platform Free, open-source pytest-style assert_test(), so agent evals run in your existing test suite

If your agent connects to external tools through MCP as part of what you're evaluating, our MCP troubleshooting guide covers the connection failures that tend to show up as false negatives in eval runs. And once evals are running in CI, pair them with the AI Agent Observability Tools that catch what evals miss in production.

The Most Common Mistake: Trusting a Benchmark You Didn't Audit

The SWE-Bench Pro retraction above isn't an isolated case — it's a pattern worth internalizing. Public benchmarks get gamed, go stale, or turn out to have been broken from the start, and reward hacking research this year found it's swamping genuine model-intelligence gains on some coding evals. Before you adopt any third-party benchmark as part of your own eval suite, spend an hour actually reading a sample of its tasks rather than trusting its reputation.

For a broader look at how coding agents are actually being chosen and compared this year, see Claude Code vs Cursor vs OpenCode.

FAQ

1. What's the difference between an eval and a benchmark?

A benchmark is a fixed public task set used to compare models against each other. An eval is something you build to test your own agent against your own task and data. You cite benchmarks; you build evals.

2. Do I need a paid platform to write agent evals?

No. DeepEval, Arize Phoenix, and Promptfoo are all free and open-source, and are fully functional at zero spend. Paid platforms like LangSmith and Braintrust add hosted collaboration, dashboards, and compliance certifications on top.

3. Should I start with final-answer evals or trajectory evals?

Start with final-answer evals — they're faster to build and give you an immediate baseline. Add trajectory evals as soon as your agent takes more than one or two steps, since that's where most real failures happen.

4. Is LLM-as-judge scoring reliable?

It's useful for fuzzy tasks a deterministic check can't handle, but it's slow, costly at scale, and can be gamed. Score outcomes that are hard to satisfy superficially — like Cursor's post-merge resolution check — rather than surface-level plausibility.

5. Can I trust a well-known benchmark like SWE-Bench without checking it myself?

Not automatically. OpenAI's own July 2026 audit found roughly 30% of SWE-Bench Pro's tasks were broken and retracted its recommendation to use it. Audit a sample of any benchmark's tasks before building your evaluation strategy around it.

Author Image

Hardeep Singh

Hardeep Singh is a tech and money-blogging enthusiast, sharing guides on earning apps, affiliate programs, online business tips, AI tools, SEO, and blogging tutorials. About Author.