Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 2 min read

Can AI Agents Actually Use Your Website? I Built a Tool to Find Out

AI agents are getting better at using websites, but there is still a basic question that is surprisingly difficult to answer: Can an AI agent actually use your website? I built Agentic Web Check to test that question i

AI agents are getting better at using websites, but there is still a basic question that is surprisingly difficult to answer:

Can an AI agent actually use your website?

I built Agentic Web Check to test that question in a real browser.

The idea is deliberately simple: instead of only checking whether a website has the right HTML, metadata, accessibility signals, or AI-facing interfaces, the tool actually opens the site in Chromium and tries to use it.

It has three main parts:

  • deterministic checks of what an agent can perceive and interact with
  • real browser tasks such as finding contact information, privacy policies, or help pages
  • programmatic verification of whether each task actually succeeded

The last part is the one I care about most.

The agent does not get to say β€œdone” and receive a PASS.

A task ends as PASS, FAIL, BLOCKED, or INCONCLUSIVE based on assertions against the resulting browser state.

There is also a separate safety layer that looks for things such as hidden instructions aimed at AI systems, unsafe forms, authentication boundaries, downloads, and consequential actions without appropriate safeguards.

One of the controlled fixtures gave an interesting result: it scored 81 on the perception checks, but all three browser tasks ended BLOCKED because its consent banner did not expose a properly named close control.

That is a deliberately broken fixture, not a claim about the wider web. But it illustrates the distinction I wanted to measure: a page can look reasonably good to a static audit while still being difficult for an agent to operate.

The project is still very early.

It is currently v0.1.0, Chromium-focused, does not perform a full site crawl, and I have not published a public benchmark. Some thresholds are documented defaults rather than empirically validated constants.

The goal is to make those limitations visible and improve the methodology through real use rather than hiding them behind a single β€œAI readiness” score.

The project is open source:

https://github.com/ericovirgy/agentic-web-check

I would be particularly interested in feedback from people working on browser agents, Playwright, accessibility, WebMCP, or agentic web tooling.

What would you measure differently?

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.