Dev.to AI 🤖 Ai 👁 0 📖 2 min read

Write the Acceptance Tests Before the Agency Writes the Automation

If you are the technical person asked to "work with the agency", here is the one thing I would do before kickoff: write the acceptance tests yourself, in plain language, and attach them to the contract. It changes the co

If you are the technical person asked to "work with the agency", here is the one thing I would do before kickoff: write the acceptance tests yourself, in plain language, and attach them to the contract. It changes the conversation with an AI automation agency that writes acceptance tests first or with any vendor, because "done" stops being a feeling.

Disclosure: I run NxFlowAI. We ask clients for exactly this, but the format works with anyone.

Why acceptance tests, not requirements

Requirements describe features. Acceptance tests describe behaviour with real inputs. For AI workflows that matters, because the same feature ("reply to enquiries") can behave very differently on a voice note, a duplicate lead or an angry customer.

A small format that works

- id: AT-01
  given: "New WhatsApp message from an unknown number asking for opening hours"
  expect: "Auto-reply from the approved library within the business's reply window; lead created in CRM with source=whatsapp"
  human_approval: false
- id: AT-02
  given: "Existing customer asks for a discount on a quoted order"
  expect: "Draft reply created, NOT sent; owner notified; CRM deal unchanged"
  human_approval: true
- id: AT-03
  given: "Same person messages on WhatsApp and fills the web form within 10 minutes"
  expect: "One CRM contact, two activities, no duplicate deal"
  human_approval: false
- id: AT-04
  given: "CRM API returns an error"
  expect: "Message queued, retried, alert sent to the named owner after the retry limit"
  human_approval: false

Rules of thumb

  1. Take inputs from real messages. Export a sample, remove personal data, and use those as the given lines.
  2. Write at least as many failure tests as happy-path tests. Errors, duplicates, opt-outs, after-hours.
  3. Mark every test that must involve a person. If a test sends money, prices or apologies, human_approval should be true.
  4. Agree who runs the tests. The agency runs them before go-live; you rerun a sample after every change.
  5. Keep them in the repo or the shared drive, next to the workflow map.

What this does to the proposal

Proposals get shorter and more honest. Vague items like "smart routing" turn into AT-05 to AT-09. Price conversations get easier too, because you can point to the tests that drive the effort.

Our first step is a 72-hour audit of one workflow, and these tests come out of it. If you are still deciding whether you need custom work at all, this comparison of custom AI and off-the-shelf chatbots is a useful filter.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.