How to Write Evals for AI Agents: A Step-by-Step Guide
Most teams that ship an AI agent test it the way they'd test a chatbot: run some prompts, eyeball the answers, ship it. That works until the agent starts calling tools, taking multiple steps, and failing in ways a single right-or-wrong answer can't catch — looping on a step, calling the right tool with the wrong arguments, or technically completing the task in a way that breaks something else. Writing real evals is how you catch that before your users do. Here's how to actually build them.
Step 1: Know the Difference Between a Benchmark and an Eval
A benchmark is a fixed, public
task set with a leaderboard — it tells you how a model compares to other models
on a standardized problem. An eval is something you build to test your own
agent against your own task, on your own data. You cite a benchmark; you build
a framework. If your only testing plan is “check how it does on a public
benchmark,” you're comparing models, not verifying your agent actually works.
This distinction matters more
than it sounds. In July 2026, OpenAI audited SWE-Bench Pro — a widely cited
coding benchmark — and found that roughly 30% of its 731 tasks were broken:
overly strict tests that rejected correct solutions, underspecified prompts,
and tasks that passed on incomplete fixes. OpenAI retracted its own earlier
recommendation to use it. If a widely respected benchmark can be a third
broken, your own evals — built around your actual task — matter more than
borrowing someone else's leaderboard.
Step 2: Decide Which Layer You're Actually Testing
Agent evaluation has three
layers, and most teams only build the first one:
•
Final-answer evaluation — score the last message
against an expected result. Necessary, but the answer can be right while the
path to it was slow, expensive, or unsafe.
•
Trajectory evaluation — score the sequence of steps
that produced the answer: did it call the right tools with the right arguments,
and did it loop or recover cleanly from a bad tool call?
•
Online evaluation — score live production traffic in
real time, using a fast classifier rather than a slow, expensive LLM-as-judge
call on every single turn.
Start with final-answer evals
because they're the easiest to build, but budget for trajectory evals before
you ship anything that takes more than one or two steps — that's where most
real agent failures actually live.
Step 3: Build a Golden Dataset From Real Traces, Not Just Synthetic Cases
A golden dataset is the set of
representative tasks you'll score your agent against every time you change a
prompt, a tool, or a model. Synthetic test cases are a fine starting point, but
the highest-value entries come from real production traces — actual
conversations where your agent looped, called the wrong tool, or frustrated a
user. Pull those into your dataset as regression tests so a fixed bug can't
silently come back.
If you're building this
alongside a coding agent specifically, our AI
Coding Agents Explained guide and Best
AI Agent Framework for Python cover the frameworks these datasets typically
plug into.
Step 4: Pick a Scorer — and Know Where It Breaks
Deterministic checks (file
exists, regex match, exact string comparison) are fast and cheap but only work
for narrowly defined tasks. LLM-as-judge scoring handles fuzzier tasks but is
slow, expensive at scale, and vulnerable to reward hacking — where an agent (or
a model being trained against your eval) learns to satisfy the scorer's letter
rather than its intent.
Cursor's own writeup on its
Bugbot code-review tool is a useful real-world example: instead of scoring
whether the agent's suggestion looked plausible, they built their primary
metric around post-merge signal — using a separate AI check at merge time to confirm
a flagged bug was actually fixed by the developer, validated with human
spot-checks. That single change, plus 40 subsequent experiments, moved their
resolution rate from 52% to over 70%. The lesson generalizes: score outcomes
that are hard to game, not surface plausibility.
Step 5: Wire Evals Into CI/CD So Regressions Get Caught Automatically
An eval you run manually before a big release is better than nothing, but an eval that runs on every pull request is what actually prevents regressions. Most of the tools below support this directly — either as a native CI integration or, in DeepEval's case, as a pytest assertion that runs alongside your existing test suite.
| Tool | Best For | Pricing | Standout Detail |
|---|---|---|---|
| LangSmith | Teams already on LangChain/LangGraph | Free (5k traces/mo); Plus $39/seat/mo | SOC 2 Type II at its $39 tier — the lowest compliance entry point in the category |
| Braintrust | Product teams shipping fast, dataset-first workflows | Free Starter; Pro $249/mo flat-rate | Sandboxed custom Python scorers no other platform currently offers |
| Arize Phoenix | Open-source production tracing + eval in one stack | Free, self-hosted (Elastic License) | OpenTelemetry-native; parent company Arize was announced for acquisition by Dynatrace in August 2026 |
| Promptfoo | Security and red-team testing of agents | Free, open-source | 500+ built-in attack vectors; acquired by OpenAI in March 2026 |
| DeepEval | CI-integrated testing without a hosted platform | Free, open-source | pytest-style assert_test(), so agent evals run in your existing test suite |
If your agent connects to
external tools through MCP as part of what you're evaluating, our MCP
troubleshooting guide covers the connection failures that tend to show up
as false negatives in eval runs. And once evals are running in CI, pair them
with the AI
Agent Observability Tools that catch what evals miss in production.
The Most Common Mistake: Trusting a Benchmark You Didn't Audit
The SWE-Bench Pro retraction
above isn't an isolated case — it's a pattern worth internalizing. Public
benchmarks get gamed, go stale, or turn out to have been broken from the start,
and reward hacking research this year found it's swamping genuine model-intelligence
gains on some coding evals. Before you adopt any third-party benchmark as part
of your own eval suite, spend an hour actually reading a sample of its tasks
rather than trusting its reputation.
For a broader look at how
coding agents are actually being chosen and compared this year, see Claude
Code vs Cursor vs OpenCode.
FAQ
1. What's the difference between an eval and a
benchmark?
A benchmark is a fixed public
task set used to compare models against each other. An eval is something you
build to test your own agent against your own task and data. You cite
benchmarks; you build evals.
2. Do I need a paid platform to write agent evals?
No. DeepEval, Arize Phoenix,
and Promptfoo are all free and open-source, and are fully functional at zero
spend. Paid platforms like LangSmith and Braintrust add hosted collaboration,
dashboards, and compliance certifications on top.
3. Should I start with final-answer evals or trajectory
evals?
Start with final-answer evals —
they're faster to build and give you an immediate baseline. Add trajectory
evals as soon as your agent takes more than one or two steps, since that's
where most real failures happen.
4. Is LLM-as-judge scoring reliable?
It's useful for fuzzy tasks a
deterministic check can't handle, but it's slow, costly at scale, and can be
gamed. Score outcomes that are hard to satisfy superficially — like Cursor's
post-merge resolution check — rather than surface-level plausibility.
5. Can I trust a well-known benchmark like SWE-Bench
without checking it myself?
Not automatically. OpenAI's own
July 2026 audit found roughly 30% of SWE-Bench Pro's tasks were broken and
retracted its recommendation to use it. Audit a sample of any benchmark's tasks
before building your evaluation strategy around it.
Hardeep Singh
Hardeep Singh is a tech and money-blogging enthusiast, sharing guides on earning apps, affiliate programs, online business tips, AI tools, SEO, and blogging tutorials. About Author.

Comments
Post a Comment