When you change an AI agent’s prompt, how do you check that you fixed one problem without creating another?
A prompt change fixes one issue, but how do you check it hasn’t caused another?
Do you rerun saved test cases or check a few conversations manually? Curious what’s worked for your team.