Skip to content
AG-02Building AI agents7 checks · 7 min

Where agents actually fail once real people use them

Read this before you sign off a pilot. Every item here is something that works in testing and breaks in production.

  1. 01

    It is confidently wrong rather than unsure.

    A model's fluency is unrelated to its accuracy, so the wrong answers arrive in the same calm tone as the right ones. Users learn to trust it during the easy weeks and then get caught. The fix is citations and a visible confidence boundary, designed in from the start.

  2. 02

    The demo questions were the easy ones.

    Whoever built the demo unconsciously asked questions it could answer. Real users arrive with two questions at once, a typo, an attachment, and context from an email you cannot see.

  3. 03

    It cannot say 'I don't know'.

    in the PDF

  4. 04

    Someone finds the edge of the guardrails and pushes.

    in the PDF

  5. 05

    It succeeds and nobody notices the side effect.

    in the PDF

  6. 06

    Latency makes it useless even though it is correct.

    in the PDF

  7. 07

    It quietly degrades after a provider update.

    in the PDF

5 more checks in the PDF

The complete checklist — every check with the reasoning under it, the closing note, and the other 12 documents — comes as one PDF.

One email, one PDF, all 13 documents. We don't run a newsletter — unless a new one goes up, that's the only message you'll get.

Before you build

13 numbered documents on scoping software, building it, and taking delivery of it. Free, and free to copy — put them in your own wiki, take them into your own meetings, send them to whoever is about to sign something.

Free to use, share and adapt. No attribution required. © 2026 TechBuilds.