Build Agents That Keep Their Rules: What to Test Before You Ship a Vertical Agent
Subtitle: Multi-turn drift, underspecified prompts and a blank-box UX are the three failure points to design around. By Tej Pandya, founder of GrowEasy.ai I think vertical agents will do well because they fix three pro
Subtitle: Multi-turn drift, underspecified prompts and a blank-box UX are the three failure points to design around.
By Tej Pandya, founder of GrowEasy.ai
I think vertical agents will do well because they fix three problems a general agent leaves to the user. If you build one, these are the things to test.
Test that rules survive later runs, not just the first
In my own work, agreed rules slip a few deliveries later. That is my observation, not a measured rate. The ICLR 2026 paper "LLMs Get Lost in Multi-Turn Conversation" reports an average 39% drop across six generation tasks when instructions arrive step by step (15 models, 200,000+ simulated conversations). The ACL Findings 2026 paper "What Prompts Don't Say" reports underspecified prompts are about 2x as likely to regress across model or prompt changes. Both are lab tests of underspecified instructions. Build a regression check that replays your rules after many runs and after every model or prompt change.
Write the long prompt once and version it
DETAIL (Kim, Dec 2025; 30 tasks, GPT-4 and o3-mini) found specificity improved accuracy, most for smaller models and procedural tasks. Weak evidence, but it matches practice. State the goal, steps, limits and quality bar, and ship them with the agent, so users do not write them.
Start users at a result
A CMU study (31 participants, Operator and Manus) found usability barriers with general agents, including capabilities that did not match user expectations. It does not say people lack ideas. My opinion: a blank box asks users to invent the job, so ship an agent for one job.
Know the risks
Boom forecasts come from investors and analysts. Thin wrappers can be absorbed by model vendors. Narrow agents still need rule checks. For legal, health and finance work, keep a human sign-off.
Sources: ICLR 2026, ACL Findings 2026, CMU study, DETAIL.
Video version: https://youtu.be/bGr6M1ddN_U
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.