Skip to content
AG-03Building AI agents5 prompts · 6 min

Agent eval set starter

Run these in order against your own tickets. The output is an eval set, which is the thing that tells you whether a prompt change helped or just moved the failure.

  1. 1 · Mine real cases from the inbox

    You have an export of past tickets and need the twenty that matter.

    Here are [N] support messages from our inbox.
    
    [PASTE TICKETS]
    
    Group them by the underlying request, not by wording. For each group give me: the request in five words, how many messages fall into it, and one verbatim example that is representative rather than the clearest. Order by frequency. Flag any group where the messages disagree about what the customer wanted — those are the ones worth looking at twice.
  2. 2 · Turn one ticket into an eval case

    For each group from step 1.

    in the PDF

  3. 3 · Write the grader

    Once you have cases and need to score them without reading every one.

    in the PDF

  4. 4 · Find the cases you have not thought of

    Before launch. This is the step people skip.

    in the PDF

  5. 5 · Turn an incident into a regression test

    Every time it gets something wrong in production.

    in the PDF

4 more prompts in the PDF

The complete pack — every prompt, the closing note, and the other 12 documents — comes as one PDF.

One email, one PDF, all 13 documents. We don't run a newsletter — unless a new one goes up, that's the only message you'll get.

Before you build

13 numbered documents on scoping software, building it, and taking delivery of it. Free, and free to copy — put them in your own wiki, take them into your own meetings, send them to whoever is about to sign something.

Free to use, share and adapt. No attribution required. © 2026 TechBuilds.