We asked an AI agent the same question ten times. Then we changed one block of the site.
I am Florin Livada. I run Livada (livada.io), a small company that sells SEO and AI-visibility tools. What follows is a measurement on our own test site, not a customer case, and it is not an advert. AI assistants answe
I am Florin Livada. I run Livada (livada.io), a small company that sells SEO and AI-visibility tools. What follows is a measurement on our own test site, not a customer case, and it is not an advert.
AI assistants answer questions about your site. Do their answers contain the facts a visitor needs? You can measure that instead of guessing.
We tested it on our own pilot site, liveverdon.com, a local guide to the Verdon gorges in France (not a customer). One real question: βHow long does it take to hike the Imbut trail in the Verdon gorges, and is it difficult?β That trail has been closed since 5 July 2022 by a municipal order. The trail's own page said so. The home page, where the agent starts, did not.
Method, in six steps:
- One real question, not an invented query.
- An AI agent (OpenAI's gpt-5.4-mini) in a real browser gets the question ten times, starting from the home page of the site as it is.
- We fix one block of the site. The site's content fingerprint is recorded on every run and is identical within each series of ten.
- The same agent gets the same question ten times again.
- Each answer is judged fact by fact, against four facts written down in advance, by a model from a different provider than the agent (Anthropic's Claude Haiku 4.5), so a model never grades itself.
- Raw results go into a signed, chained, dated log that is never rewritten.
The four facts: the trail is closed; the closure is still in force; the 4 to 6 hours figure is from when it was open; it is difficult to very difficult. An answer counts only if it carries all four.
Result, 3 September 2026: before the change, 0 of 10 answers carried the four facts. All ten said the agent could not find information about the trail on the page it had. After the change, 10 of 10 did: the trail is closed at present, and the 4 to 6 hours dates from when it was open. Ten passes is a small number: the statistical range for β10 of 10β runs from about 7 to 10 out of 10, so read it as βworked on this wordingβ, not βalways worksβ.
What this does not show: one site, one question, one agent model, one judge model, at the dates of the tests. We tried several wordings of the fix with the same agent beforehand and kept one that worked, so this is an exploratory result, not a confirmation. The ten passes were nearly identical (two distinct answers out of ten after the change): it shows repeatability on one wording, not robustness to other wordings. Nothing here predicts what an agent will do tomorrow. An earlier attempt on 2 September produced a false pass and we withdrew it. In two other cases we published, the headline figure comes from a confirmation series run after we widened the question, so it is not a strict before/after; the page says so.
The protocol and the published measurements are described at: https://livada.io/en/research/?utm_source=devto&utm_medium=article&utm_campaign=devto-0610
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.