15 July 2026 · 1 min read
The difference between an agent that survives contact with real customers and one that gets switched off in week three is usually 200 rows in a spreadsheet.
It is a list of real inputs your agent will see, each paired with what a good response looks like. Fifty is a minimum. Between one and two hundred is where most of the value sits.
It is not synthetic. If you generate the examples with a model, you are testing the model against its own assumptions, which is how agents pass every test and then fail on the first customer who types in Arabic or attaches a photo of a receipt.
Your existing records. Closed tickets, past enquiries, processed invoices, resolved matters. The awkward ones matter more than the clean ones, so oversample the escalations and the complaints.
Autonomy is not a setting. It is something an agent earns, one evaluation run at a time.
You do. The evaluation set is the most valuable artefact of an agent project, more valuable than the prompt, because it is the thing that lets anyone maintain the agent later. It should be in your repository, in your name, from the first week.
Written by the OneRee team. If any of this is wrong for your situation, tell us — we would rather update the page than be right on the internet.
Keep reading
Next step
Thirty minutes, no deck. We will tell you what we would do first, what it costs, and whether we are the right people for it.