@ankrgyl
Post
Post 1 of 2
I have long felt the diminishing effectiveness of LLM-as-a-judge for agents, and one day over coffee @mitch\_troy gave me a rant that helped clarify why. The ground truth for an agent isn't its output, it is how the agent should behave. This is important for two reasons: (1) it is very difficult to generate good ground truth values (2) to effectively debug why an agent produced invalid outputs, you need to introspect its behavior. Mitch then walked through why this is so hard for their team's tax agents. The behaviors themselves are subtle, hard-earned lessons from analyzing lots of failures in their traces in Braintrust, and to effectively flag them, you must effectively document them. That discussion led us to build behavior specs: a new open standard we're releasing with @trybasis that documents how agents should behave. The repo includes a definition of the spec along with open examples for how to write good specs, evaluate agents using it, and more. Clear, human-articulated prose is the highest leverage way to drive agentic systems to produce great outcomes. Behavior specs provide a framework to do that with evals. I'm super excited to work on this in the open, and would love to get feedback from others on how we can make this spec maximally useful. Please try it out, and share your thoughts! https://www.agentbehavior.dev/ Quoted post by Mitchell Troyanovsky (@mitch\_troy) Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days.
Even if you had a reliable way to verify outcomes at scale (and weren’t bothered by the multi-hour iteration loops), the sheer volume of decisions by the agent that occur in a multi-hour job makes it hard to know whether performing well will generalize to production.
Over the last two years at @trybasis, we've been solving this problem by supervising the process our agents take to get to outcomes, rather than just looking at whether the outcome itself is correct.
We think this is the key to building production agents at scale.
It's what has allowed us to run agents in production that operate for hours, sometimes days, and reliably perform tasks like entire complex tax returns end to end.
Today, alongside @braintrust, we're open sourcing a standard for defining, evaluating, and eventually rewarding agent behaviors.
Thread below with all the details on how we’re scaling behaviors to close the loop for long-horizon agents.
Image from X post Open quoted post on X
Post 2 of 2
If you want to see how behavior specs work with evals in Braintrust, @daRubberDuckiee made a great demo video:
Video thumbnail from X post Watch video
Explanation
What it says For long-horizon agents, “did the final answer pass?” is an inadequate eval. An agent may make hundreds of consequential intermediate decisions over hours or days, so the proposed ground truth is the behavior/process it should follow: how it gathers context, decides, acts, handles uncertainty, and recovers. Braintrust and Basis are open-sourcing “behavior specs” as a standard for writing those expectations down and evaluating traces against them.
Context The idea came from Basis’s experience building tax agents, where outcomes are slow, difficult to verify, and sparse as training/eval signals. Their claim is that supervising process—not merely final outputs—has enabled agents to execute complex tax work for hours or days more reliably.
Why it matters This reframes agent evals from “LLM judge scores the answer” toward something closer to executable operating procedures plus trace-based auditing. The interesting claim is that explicit human-written behavioral knowledge can become a shared source of truth connecting prompts, reviews, evals, and eventually rewards. Whether this actually scales without specs becoming enormous, brittle, or gameable is the key unresolved question.
Images The spec is concretely a `BEHAVIOR.md` file with YAML frontmatter plus free-form Markdown, stored under `.agents/behaviors/`; optional references can contain rationale, examples, and background material.