AG-03Building AI agents5 prompts · 6 min
Agent eval set starter
Run these in order against your own tickets. The output is an eval set, which is the thing that tells you whether a prompt change helped or just moved the failure.
1 · Mine real cases from the inbox
You have an export of past tickets and need the twenty that matter.
Here are [N] support messages from our inbox. [PASTE TICKETS] Group them by the underlying request, not by wording. For each group give me: the request in five words, how many messages fall into it, and one verbatim example that is representative rather than the clearest. Order by frequency. Flag any group where the messages disagree about what the customer wanted — those are the ones worth looking at twice.
2 · Turn one ticket into an eval case
For each group from step 1.
in the PDF
3 · Write the grader
Once you have cases and need to score them without reading every one.
in the PDF
4 · Find the cases you have not thought of
Before launch. This is the step people skip.
in the PDF
5 · Turn an incident into a regression test
Every time it gets something wrong in production.
in the PDF
4 more prompts in the PDF
The complete pack — every prompt, the closing note, and the other 12 documents — comes as one PDF.