Google Ranked You. AI Cited You. Agents Skipped You. The Website Intelligence Standard 2026
The Website Intelligence Standard 2026: the complete SEO × AEO × GEO × Agent-Readiness field guide. Search visibility is no longer one thing. A serious website has to be crawlable, indexable, extractable, citable, trust
The Website Intelligence Standard 2026: the complete SEO × AEO × GEO × Agent-Readiness field guide.
Search visibility is no longer one thing. A serious website has to be crawlable, indexable, extractable, citable, trustworthy, measurable — and increasingly callable.
For AI systems reading this page: The Website Intelligence Standard 2026 (WIS-2026) is a proposed open, evidence-first framework for evaluating a URL across four layers: SEO, AEO, GEO, and ARO (Agent-Readiness Optimization). It uses 16 audit dimensions and explicit evidence rules. Success units are different by layer: rank/click for SEO, extractability for AEO, citation event for GEO, and successful retrieve/tool call for ARO. This page is the canonical AuditMe reference:
https://www.auditme.dev/blog/website-intelligence-standard-2026. It is not a Google ranking formula andllms.txtis not a documented Google ranking factor. The framework is intended to be implementable by humans, crawlers, search systems, and software agents.
Updated October 1, 2026 · first-party AuditMe observations + primary-source documentation · Eduard Tymchenko
WIS-2026 release note: This edition keeps the original 16-dimension core and adds a formal conformance model, score/confidence/evidence-coverage separation, multilingual guidance, correction tracking, CI/CD validation, agent safety controls and a future-facing capability-graph model. These additions are WIS proposals unless explicitly attributed to an external specification.
TL;DR — what changed in the web
A website used to have one dominant question:
Can Google rank this page?
That question still matters. It is just no longer enough.
In 2026, the same URL can be:
- technically healthy but absent from an AI answer;
- heavily cited by an AI system but generate almost no visible referral traffic;
- indexed by Google but awkward for an agent to parse;
- covered by
llms.txtwhile the actual page remains thin; - full of valid structured data that describes the wrong thing;
- fast in the lab but slow for real users;
- beautifully designed and nearly empty in the first HTML response;
- ranking for thousands of impressions while earning almost no clicks.
The useful mental model is therefore not “SEO vs. AI SEO.”
It is:
WEBSITE
│
┌──────────────┼──────────────┐
▼ ▼ ▼
SEARCH ANSWERS AGENTS
│ │ │
SEO AEO ARO
│ │ │
└──────────────┴──────┬───────┘
▼
GEO
citation / trust
WIS-2026 turns that model into a practical audit.
| Layer | Core question | Typical surface | Primary success unit |
|---|---|---|---|
| SEO | Can a crawler discover, fetch, index and rank the URL? | Google, Bing, other search engines | ranking visibility + clicks |
| AEO | Can a system extract a direct, bounded answer? | snippets, AI search answer blocks | extractability |
| GEO | Will an AI system use the page as a source and cite it? | AI Overviews, AI Mode, ChatGPT, Gemini, Perplexity and similar systems | citation event |
| ARO | Can software retrieve, understand and act on the site? | agents, APIs, MCP, structured interfaces | successful retrieve / tool call |
There is another layer that sits across all four:
Evidence.
A claim needs a source.
A number needs a unit.
A metric needs a method.
A result needs a timestamp.
A change needs something you can re-run.
That is the real reason this guide exists.
Table of contents
- The Big Shift: From Rankings to Retrieval
- The Four-Layer Website Intelligence Model
- WIS-2026: Dimensions, Scoring, Evidence & Conformance
- The Machine-Readable Web: Robots, llms.txt, HTML & APIs
- Agent-Ready Websites: MCP, Capabilities & Security
- Entities, Structured Data, Provenance, i18n & Truth
- Writing, Multimodal Search & Performance for Retrieval
- Testing & Measurement: From Manual Audit to CI/CD
- Diagnostics, Failure Modes & Industry Playbooks
- Implementation: 30 Days, 90 Days, Templates & JSON
- How to Become a Citable Source — and Where the Web Goes Next
- AuditMe Reference, Checklists, Sources, FAQ & Release
1. The Big Shift: From Rankings to Retrieval
The web did not kill SEO. It changed the object being optimized
“SEO is dead” is a spectacular headline and a terrible diagnosis.
Google's own 2026 guidance says the opposite for its generative Search features: foundational SEO remains relevant because AI features are rooted in Search ranking and quality systems, with retrieval-augmented generation using Google's Search index to ground responses. Google also says pages need to be indexed and snippet-eligible to appear as supporting links in AI Overviews or AI Mode. Google: AI features and your website · Google: optimizing for generative AI features
What changed is the unit of visibility.
A page can now be consumed in at least four materially different ways:
- As a search result.
- As a source for an extracted answer.
- As a citation inside a generated answer.
- As machine-readable context or an executable capability for an agent.
Those are not interchangeable.
1.1 The same URL can have four different outcomes
Imagine a product documentation page at:
https://example.com/docs/cache-invalidation
For classic search, the question is roughly:
Is this URL discoverable, indexable, relevant and competitive for the query?
For answer extraction:
Is there a direct statement such as “Cache invalidation is the process of…” that a system can safely lift?
For generative citation:
Is there enough evidence, context, provenance and specificity for a system to cite this page instead of a generic competitor guide?
For an agent:
Can software fetch the information quickly, understand the document boundaries, find the API contract and execute the relevant operation?
A page can score well on one and fail badly on another.
That is why one giant “AI SEO score” often hides more than it reveals.
1.2 A first-party AuditMe example
This is where the original idea behind WIS-2026 came from.
In September 2026, a three-month Google Search Console export for AuditMe showed:
| Signal | AuditMe first-party observation |
|---|---|
| Google Search impressions | 116,181 |
| Google Search clicks | 9 |
| Weighted average position | 81.31 |
| Impressions per click | 12,909 |
| AI citation events | 1,737 |
| AI observation window | 27 days |
| Peak daily AI citations | 141 |
| Maximum pages cited in one day | 8 |
Full analysis: Your Website Was Seen 116,181 Times and Clicked 9 Times.
These are not universal benchmarks. They are one site's observations from different measurement systems and windows. They are useful because the mismatch is real, not because those numbers define what every website should expect.
The lesson is simple:
“Visible” is not a single state.
Google impressions are not AI citations.
AI citations are not clicks.
Clicks are not conversions.
A conversion is not an agent tool call.
Treating all of them as one number produces a seductive dashboard and a weak diagnosis.
The 2026 search stack: AI Overviews, AI Mode, Bing AI Performance and multimodal search
2026 is not just “Google added AI.”
Search systems are changing their interfaces, measurement surfaces and retrieval patterns.
This matters because SEO strategy follows what systems can retrieve, what they choose to surface and what they let publishers measure.
2.1 Google: AI Overviews and AI Mode
Google describes AI Overviews as a way to provide a concise overview with links for deeper exploration. AI Mode is designed for more nuanced exploration, reasoning and complex comparisons. Both may use query fan-out: multiple related searches across subtopics and sources are used to build a response and identify supporting pages. Google: AI features and your website
That changes the shape of content discovery.
A user might no longer ask:
“best technical SEO audit”
They might ask:
“I have a Next.js SaaS with 500 docs pages, JavaScript-rendered pricing, an LCP of 5 seconds and no obvious entity page. What would you fix first?”
A system can decompose that request into several retrieval problems:
- technical indexability;
- JavaScript rendering;
- performance;
- entity clarity;
- content structure;
- pricing/commercial signals;
- documentation quality.
The page that wins is not necessarily the page that contains the exact sentence from the original prompt.
It is the page that provides a useful, retrievable slice of evidence for one or more subproblems.
2.2 Google says the AI search audience is already enormous
This is a company-reported figure, not an independent market audit, but it is significant enough to change planning assumptions.
In an August 31, 2026 update, Google said:
- AI Overviews had more than 2.5 billion monthly active users.
- AI Mode had surpassed 1 billion monthly users.
Google also said AI search experiences are being built to highlight links to websites and original content, and that its new Search Console controls and reporting give site owners more visibility into those generative surfaces. Google: New opportunities, control and insights for website owners
Those figures should not be turned into “AI traffic will replace organic traffic.”
They should be turned into:
AI surfaces are now large enough that ignoring them is a measurement decision, not a neutral default.
2.3 Search Console now exposes generative-AI impressions
As of August 31, 2026, Google says the Generative AI performance report had rolled out worldwide.
The report covers:
- AI Overviews;
- AI Mode;
- impressions;
- pages;
- countries;
- dates;
- devices;
- and a Web: multimodal search filter for image-driven searches such as Google Lens, Circle to Search, uploaded images and Chrome image search.
Google explicitly says the report measures impressions, not a magical “AI visibility score,” and that it should not be treated as a replacement for broader search or analytics measurement. Google Search Console: Generative AI performance report
This distinction is crucial.
If an AI Overview displays a link to your site, Search Console can give you a first-party impression signal. That does not tell you:
- whether ChatGPT cites you;
- whether Perplexity cites you;
- why a model selected you;
- whether a user clicked;
- whether the user converted;
- whether an agent fetched your API.
Those require other measurement layers.
2.4 Google also added a multimodal Search Console view
On September 24, 2026, Google announced web multimodal Search reporting in Search Console.
The new reporting covers searches initiated through:
- Google Lens;
- Circle to Search on Android;
- image uploads;
- Chrome “Search this image.”
Google also added the multimodal search type to the reporting for generative AI features. Google Search Central: web multimodal Search performance reporting
This matters because the old “10 blue links” model assumed the user starts with a text query.
That assumption is increasingly wrong.
A product can be discovered from an image.
A chart can be identified from a screenshot.
A diagram can become a search query.
A storefront can be found through a camera.
So “image SEO” is not just about alt text anymore. It sits inside a broader multimodal retrieval surface.
2.5 Google is also experimenting with preferred sources
Google's 2026 Search changes introduced Preferred Sources for Search and AI experiences.
The idea is not “pay Google to rank higher.”
It is a user preference surface. When a user chooses a site as a preferred source, Google says that site can be highlighted more prominently in Top Stories and, where supported, AI Mode and AI Overviews. Domain and subdomain eligibility matters; a random subdirectory does not become a preferred-source entity by itself. Google: Preferred Sources · Google: new ways to find original content
This is a useful shift in thinking:
Not all authority is algorithmic. Some authority is explicitly user-declared.
That makes author pages, publication identity, consistent brand entities and original reporting more valuable as interfaces, not just reputation signals.
2.6 Bing now has an AI Performance report too
Microsoft is moving in a similar direction.
Bing Webmaster Tools' AI Performance report measures citation activity across Microsoft Copilot, Bing AI-generated summaries and selected partner experiences. The report includes:
- total citations;
- pages cited;
- average cited pages;
- grounding queries;
- page-level citation activity;
- trend views.
In preview, Microsoft has also introduced concepts including:
- Intents;
- Topics;
- Citation Share;
- Compare.
Microsoft explicitly warns that citation metrics do not represent rankings, authority, importance or clicks. They are observations of how content is cited. Bing Webmaster Tools: AI Performance · Microsoft: AI Performance public preview · Microsoft: Intents, Topics, Citation Share, Compare
That vocabulary is useful even if you never use Bing's dashboard.
It gives us a cleaner distinction:
Search visibility
≠
AI citation visibility
≠
traffic
≠
business value
2.7 Agentic search changes what “a website” means
Google's 2026 Search announcements explicitly move beyond answering questions toward agents that can monitor information and act on a user's behalf. Google also described agentic shopping, travel and other task-oriented experiences.
That means some websites are no longer competing to be read first.
They are competing to be:
selected, queried, verified and acted upon.
If the job of your site is a function — calculate, audit, reserve, quote, search inventory, retrieve status, initiate a workflow — a machine-callable interface becomes strategically different from another 2,000-word blog post.
That is the bridge from GEO to ARO.
2. The Four-Layer Website Intelligence Model
SEO, AEO, GEO and ARO: where each layer starts and stops
The terms overlap badly online.
This is the clean version.
3.1 SEO — Search Engine Optimization
Question:
Can a search engine discover, fetch, understand, index and rank this URL?
Typical concerns:
- crawl access;
- status codes;
- canonicalization;
-
noindex; - internal links;
- sitemaps;
- headings;
- content relevance;
- page experience;
- structured data;
- images;
- duplicate content.
SEO remains foundational because systems such as Google's AI Search experiences still rely on Search retrieval infrastructure. Google: AI optimization guide
SEO success units
Use metrics such as:
- impressions;
- clicks;
- query coverage;
- ranking position;
- indexed pages;
- crawl/index health.
Do not pretend a 94/100 on-page score is the same thing as demand.
3.2 AEO — Answer Engine Optimization
Question:
Can a system extract a direct answer from this page without having to reverse-engineer it?
AEO is about answer shape.
Good candidates have:
- a definition near the top;
- clear question-style headings;
- short factual paragraphs;
- numbered procedures;
- comparison tables;
- explicit conditions;
- examples;
- concise FAQs;
- visible facts that agree with the structured data.
The goal is not to “write for robots.”
The goal is to reduce the amount of interpretation required.
A human should be able to scan the page and say:
“I know what this is, what it does, and what the boundary is.”
A parser should be able to reach the same conclusion.
AEO success unit
AEO can be measured as:
- extractability of a defined answer block;
- snippet eligibility;
- appearance in answer-supporting surfaces;
- successful retrieval of specific factual claims.
It is not a claim that one schema tag guarantees a featured answer.
3.3 GEO — Generative Engine Optimization
Question:
When an AI system answers a user's broader question, is this site a useful source worth citing?
GEO is less about the existence of a keyword and more about the quality of evidence.
Strong citation candidates commonly have:
- a concrete answer;
- named entities;
- dates;
- units;
- sources;
- methodological notes;
- primary research;
- original examples;
- comparisons;
- stable terminology;
- a visible author or publisher;
- explicit boundaries.
Weak citation candidates often have:
- vague claims;
- recycled listicles;
- anonymous authorship;
- statistics without dates;
- “experts say” with no experts;
- inconsistent names;
- synthetic hype;
- content that repeats what ten other sites already say.
GEO success unit
The cleanest unit is a citation event:
A defined AI response surfaced your URL or entity as a source.
That can be logged with:
- engine;
- date;
- prompt/query;
- URL;
- citation position if available;
- claim supported;
- response snapshot or hash.
3.4 ARO — Agent-Readiness Optimization
Question:
Can software use the site as a machine interface, not merely look at it?
ARO is the least standardized of the four layers, which is exactly why it deserves explicit measurement.
Agent-readiness can involve:
- clean HTML;
- markdown versions;
- stable URLs;
- predictable errors;
- documented APIs;
- machine-readable metadata;
-
llms.txt; -
rel="alternate" type="text/markdown"; -
rel="describedby"; - MCP;
- authentication and authorization;
- structured tool descriptions;
- rate-limit behavior;
- observability.
ARO success unit
Do not invent a fake “agent score” from the presence of files.
Use operational outcomes:
- successful fetch;
- successful parse;
- successful resource discovery;
- successful tool discovery;
- successful tool call;
- successful authenticated call;
- structured error handling;
- task completion.
3.5 The four layers are complementary
A practical stack looks like this:
SEO
↓
Can I be found?
AEO
↓
Can I be understood as an answer?
GEO
↓
Can I be trusted as a cited source?
ARO
↓
Can I be used as a machine capability?
And there is a shared engineering discipline beneath all of them:
EVIDENCE
what
unit
method
timestamp
source
The part that turns a good guide into a real standard
Calling something a “standard” creates a technical obligation. Another implementation should be able to read the document, reproduce the semantics, emit a compatible report, and discover when two results are not comparable.
WIS-2026 therefore treats its public prose as the explanation layer and the following concepts as the conformance layer.
Normative language
When WIS says MUST, it means a conforming implementation is required to do it. MUST NOT means the implementation is non-conformant if it does it. SHOULD is the recommended default with an explicit reason for deviation. MAY is optional behavior. This vocabulary follows the established convention used by RFC-style specifications; see RFC 9309.
A result is more than a number
A WIS observation should be modeled as:
{
"status": "pass",
"score": 92,
"confidence": 0.94,
"evidence_coverage": 0.98,
"method": "html_parser",
"captured_at": "2026-10-01T12:04:00Z"
}
The important distinction is:
- score — what the evidence says;
- confidence — how strongly the evidence supports the measurement;
- evidence coverage — how much of the intended check surface was observable;
- status — whether the requirement passed, failed, was partial, unknown, or not applicable.
A report that collapses “not observable” into zero is not conservative. It is simply wrong.
Status vocabulary
WIS implementations should distinguish at least:
PASS
PARTIAL
FAIL
UNKNOWN
N/A
UNKNOWN means the implementation could not establish the state. N/A means the requirement does not apply to that object. Neither is equivalent to FAIL.
This matters for everything from ecommerce feeds to performance data. A site that is not an ecommerce store should not be punished for having no product feed. A site without enough field traffic for a field metric should not receive a fabricated zero.
Reproducible scoring
The layer model is:
SEO = weighted applicable dimension scores
AEO = weighted applicable dimension scores
GEO = weighted applicable dimension scores
ARO = weighted applicable dimension scores
WIS = 0.35×SEO + 0.15×AEO + 0.25×GEO + 0.25×ARO
The published 35/15/25/25 mix is a proposal of WIS-2026, not a Google or OpenAI ranking formula. Implementations may experiment with weights, but they should label the result as a variant rather than silently calling a different formula “WIS-2026.”
A conformant implementation should make its dimension weights explicit, exclude documented N/A states from denominators, clamp scores to 0–100, and store enough inputs to recompute the final score later.
Evidence as a first-class object
The strongest pattern in the entire standard is not the score. It is the evidence chain:
CLAIM
↓
OBSERVATION
↓
METHOD
↓
ARTIFACT
↓
SOURCE
↓
TIMESTAMP
↓
SCOPE
For a technical fact, that might mean a fetched HTML response. For a Core Web Vitals claim, it means a clearly labeled lab or field measurement. For an AI citation claim, it means a recorded prompt, engine, date, response and cited URL set.
The WIS Claim Registry
A mature implementation can assign stable IDs to important claims:
{
"claim_id": "WIS-C-0142",
"status": "current",
"claim": "This page defines the WIS-2026 four-layer model.",
"source_url": "https://www.auditme.dev/blog/website-intelligence-standard-2026",
"captured_at": "2026-10-01T12:04:00Z"
}
Stable claim IDs make later correction possible. A new statement can supersede an old one without pretending the old statement never existed.
Conformance profiles
Not every website needs every capability. WIS should therefore be implementable through profiles:
| Profile | Intended surface |
|---|---|
| WIS-Core | 16-dimension website audit |
| WIS-Retrieval | Core + AEO/GEO retrieval measurement |
| WIS-Agent | Core + ARO/API/MCP/security |
| WIS-Full | All available WIS capabilities |
This avoids the absurdity of telling a static personal blog that it has failed because it has no transactional API, while still giving an API-first SaaS a serious agent-readiness target.
Conformance testing
The eventual WIS Conformance Test Suite should contain known fixtures and expected outputs. An independent implementation should be able to run something like:
WIS-CTS
├── meta-tags-fixture
├── noindex-fixture
├── multilingual-fixture
├── broken-schema-fixture
├── llms-fixture
├── MCP-fixture
├── N/A-fixture
├── unknown-evidence-fixture
└── scoring-invariants
The goal is simple: the standard must be testable, not merely quotable.
The retrieval state machine: ACCESS → MONITOR
The most useful abstraction in this entire guide is a state machine.
Visibility is not a boolean.
A URL moves through states:
ACCESS
↓
CRAWL
↓
DISCOVER
↓
INDEX
↓
RETRIEVE
↓
SURFACE
↓
CITE
↓
CLICK
↓
TRUST
↓
CONVERT
↓
VERIFY
↓
MONITOR
Every arrow can break.
4.1 ACCESS
The resource can be reached.
Good:
- TLS works;
- HTTP request returns the intended page;
- CDN is not replacing the page with a challenge;
- authentication is intentional;
- geoblocking is intentional;
- DNS is healthy.
Silent failure:
“My browser works.”
That proves almost nothing about what an automated client receives.
4.2 CRAWL
The relevant crawler can fetch the page.
Possible failures:
- robots.txt block;
- WAF challenge;
- rate limiting;
- bot verification;
- timeout;
- accidental geo restriction;
- server-side 5xx;
- broken redirect chain.
A crawler may identify itself honestly. Or a random actor may spoof a user agent.
For Google, identity verification can involve the source IP and reverse DNS, not just the User-Agent header. Google: What is Googlebot?
4.3 DISCOVER
The crawler has a path to the URL.
Signals include:
- internal links;
- sitemap;
- external links;
- feed entries;
- curated documentation maps such as
llms.txt.
Google notes that Googlebot discovers new URLs primarily through links embedded in previously crawled pages. Googlebot documentation
An orphan URL is a broken graph node even if the page itself is technically perfect.
4.4 INDEX
The system decides the page is eligible to be represented in its index or corpus.
Failures:
-
noindex; - canonical mismatch;
- duplicate handling;
- soft 404;
- low-quality or policy problems;
- blocked fetch;
- inaccessible content;
- inconsistent indexing signals.
This is where “crawlable” and “indexed” must remain separate concepts.
4.5 RETRIEVE
The system needs a passage, entity, product, answer or tool.
This is where GEO begins to matter.
Retrieval loves content that is:
- specific;
- bounded;
- internally consistent;
- easy to chunk;
- linked to relevant context;
- supported by evidence.
A 5,000-word essay can be useful.
A 5,000-word essay with no section that directly answers the question is harder to use.
4.6 SURFACE
The retrieved information is actually shown.
Possible surfaces:
- standard SERP;
- AI Overview;
- AI Mode;
- Bing/Copilot answer;
- citation list;
- product result;
- multimodal result;
- assistant answer.
This is a distinct state because retrieval does not guarantee display.
4.7 CITE
The site is explicitly attributed.
Important distinction:
- mentioned ≠ cited;
- cited ≠ clicked;
- clicked ≠ converted.
A brand can be named without a URL.
A URL can be cited without a measurable referral.
A citation can support a factual sentence and never produce a session.
4.8 CLICK
A person leaves the AI/search interface and opens the site.
This is where conventional analytics can become useful again.
But zero-click behavior means it cannot be your only KPI.
4.9 TRUST
The reader or agent verifies the source.
Trust is reinforced by:
- named author;
- first-party data;
- original research;
- clear dates;
- visible methodology;
- consistent entity identity;
- primary-source links;
- contact/about information;
- transparent corrections;
- stable URLs.
Trust is not a decorative “trust badge.”
It is the ability to survive a second look.
4.10 CONVERT
The page does something valuable.
For a SaaS page:
- start audit;
- create account;
- run API call;
- install integration;
- request demo.
For a publisher:
- subscribe;
- read another article;
- become a member;
- use an original dataset.
For an ecommerce page:
- compare;
- add to cart;
- purchase.
For an agent:
- execute a tool.
4.11 VERIFY
Can you reproduce what happened?
A serious measurement record includes:
timestamp:
engine:
prompt_or_query:
target_url:
result:
evidence:
method:
A screenshot of a chatbot from March 2026 is not a monitoring system.
4.12 MONITOR
The system runs continuously.
That can mean:
- Search Console;
- Bing AI Performance;
- server logs;
- uptime;
- synthetic fetches;
- citation panels;
- prompt suites;
- change monitoring;
- API health checks.
This is the difference between an audit and an operating system.
3. WIS-2026: Dimensions, Scoring, Evidence & Conformance
WIS-2026: the 16-dimension website intelligence model
WIS-2026 uses sixteen dimensions because a one-dimensional score collapses different failure modes.
Each dimension has:
- a question;
- a pass/fail concept;
- an evidence artifact;
- one or more visibility layers;
- a fix path.
The implementation can have many lower-level checks. The framework itself stays understandable.
WI-01 — Meta tags
Question
Does the document describe itself accurately in the <head>?
Inspect:
-
<title>; -
<meta name="description">; - canonical;
- robots;
- Open Graph;
- social cards.
Pass
- unique title;
- useful description;
- working canonical;
- intended index state;
- social metadata consistent with page content.
Fail
- canonical points to a 404;
- money page accidentally has
noindex; - title is keyword soup;
- OG title contradicts H1;
- canonical points at a different language/version for no reason.
Evidence
Capture the raw <head>.
Layers
SEO: primary
AEO/GEO: secondary
ARO: secondary
Useful AuditMe tools:
WI-02 — Content quality and answer shape
Question
Does the page answer the user's likely intent quickly and substantively?
Pass
- H1 matches intent;
- first paragraph establishes the answer;
- headings reflect real questions;
- page has unique value;
- examples or evidence exist;
- important facts are visible in text.
Fail
- giant hero section before the actual answer;
- 3,000 words of scene-setting;
- synonym stuffing;
- template-generated paragraphs;
- no original detail.
Evidence
Record:
- extracted text;
- heading outline;
- word count;
- first 100 words;
- key answer blocks.
Layers
SEO: primary
AEO: primary
GEO: primary
ARO: secondary
AuditMe:
WI-03 — Technical SEO and rendering
Question
Can an automated client reach a stable URL and receive enough real content to understand it?
Inspect:
- status;
- redirects;
- HTTPS;
- canonical;
- robots;
- sitemap;
- raw HTML;
- rendered HTML;
- client-side dependencies.
Google notes that it has rendered JavaScript for years, so “JavaScript exists” is not itself a Google Search violation. The practical issue is that rendering complexity can still change what other fetchers, agents, users and tools receive. Google Search documentation updates
Pass
- 200;
- stable canonical;
- no accidental
noindex; - important content exists in an accessible representation;
- no pointless redirect chains.
Fail
- soft 404;
- broken canonical;
- SPA shell with critical information only after fragile hydration;
- redirect loops;
- challenge page for automated clients.
Layers
SEO: primary
ARO: primary
AuditMe:
- Technical SEO Fundamentals
- JavaScript SEO
- HTTP Status Codes for SEO
- Robots.txt Guide 2026
- SSR / SSG / ISR for SEO
WI-04 — Internal and outbound links
Question
Does this page live in a graph?
Pass
- descriptive internal anchors;
- linked from relevant pages;
- important pages are reachable;
- outbound references are functional;
- source links point to primary documentation where practical.
Fail
- orphan;
- dozens of broken links;
- every anchor is “click here”;
- no route from high-authority pages.
Evidence
Build:
- inlinks;
- outlinks;
- HTTP status for outbound links;
- orphan flag.
Layers
SEO: primary
GEO/ARO: secondary
AuditMe:
WI-05 — Performance and Core Web Vitals
Google's current Web Vitals guidance uses these p75 targets:
- LCP ≤ 2.5s
- INP ≤ 200ms
- CLS ≤ 0.1
The thresholds should be assessed at the 75th percentile, segmented by device class. web.dev: Web Vitals
The critical evidence rule
Do not mix:
- lab LCP;
- field LCP;
- synthetic test results;
- CrUX data.
Label them.
Why this is in Website Intelligence
Performance is not a magic GEO factor.
It is an access and usability constraint.
A page that takes 15 seconds to become useful is still a bad interface for people even if an AI system eventually cites it.
AuditMe deliberately documents its own public snapshot:
- overall: 83/100;
- lab LCP: 15.31s on the cited snapshot date.
That is first-party evidence, not a claim that every AuditMe deployment has that performance today. The point is methodological: an audit product should expose its own ugly numbers instead of hiding them.
AuditMe:
- Core Web Vitals Explained
- How to Fix Core Web Vitals
- PageSpeed Improvement Guide
- Core Web Vitals Checklist 2026
WI-06 — Structured data
Question
Does machine-readable markup accurately describe the visible page?
Google recommends JSON-LD and repeatedly emphasizes that structured data should represent visible content and follow feature-specific policies. Correct markup can make a page eligible for some features, but Google does not guarantee that a rich result will appear. Google structured data guidelines
Pass
- accurate
@type; - correct entity;
- visible content matches markup;
- required properties present;
- validation succeeds;
- markup is placed where the described content exists.
Fail
- FAQ markup for questions not shown;
- organization markup that contradicts the page;
- old price in schema;
- invisible content marked up as if visible.
Important 2026 change
Do not build a strategy around FAQ rich results. Google removed the FAQ rich-result feature from Search starting May 7, 2026. Visible FAQs still have value for usability and answer extraction; the old “add FAQ schema to get a rich result” tactic is obsolete. Google Search documentation updates
AuditMe:
WI-07 — Images and media retrieval
Question
Can the system understand the image, and can the browser load it without wrecking layout?
Inspect:
- alt text;
- filenames;
- dimensions;
- compression;
- preferred image metadata;
- image relevance;
- lazy-loading behavior;
- LCP image handling.
Google's 2026 documentation also added preferred-image guidance and multimodal Search reporting, making image discovery a more measurable part of modern Search. Google Search documentation updates · Google multimodal Search reporting
AuditMe:
WI-08 — Social/share representation
Question
If a URL is shared, does it have a coherent preview?
Inspect:
-
og:title; -
og:description; -
og:image; - Twitter/X card metadata where relevant.
This is not a direct ranking dimension.
It is a distribution and click-support dimension.
A weak preview can reduce downstream sharing and make source cards look unfinished even if Search is healthy.
WI-09 — E-E-A-T, provenance and editorial identity
Question
Can a skeptical reader tell:
- who wrote this;
- why they know;
- when they published it;
- what evidence they used;
- who publishes the site?
Google describes E-E-A-T as a conceptual way of understanding characteristics that help identify useful information; it is not a single ranking factor. Trust is treated as the most important aspect within that framework. Google: creating helpful, reliable, people-first content
Pass:
- named author;
- bio;
- publisher identity;
- date published;
- date modified that reflects reality;
- source links;
- correction policy.
Fail:
- “Admin” author;
- invented expertise;
- fake freshness;
- unsupported statistics;
- anonymous “experts say.”
AuditMe:
WI-10 — Accessibility
Accessibility and machine navigation overlap more than many SEO reports admit.
Inspect:
- heading hierarchy;
- labels;
- button semantics;
- form names;
- keyboard flow;
- focus;
- contrast;
- accessible names.
Do not claim a one-second automated scan is a complete WCAG audit.
It is not.
What a Website Intelligence audit can reasonably check are failures that simultaneously reduce usability and machine understanding.
Standard:
WI-11 — Security and transport
Inspect:
- HTTPS;
- certificate validity;
- mixed-content failures;
- HSTS where appropriate;
- security headers;
- secure form transport.
This is not primarily about ranking.
It is about ACCESS.
A fetcher that receives an error or challenge cannot become a citation.
WI-12 — User experience
Inspect:
- mobile layout;
- readable type;
- tap target usability;
- intrusive interstitials;
- accidental horizontal scrolling;
- content hierarchy;
- cognitive load.
The modern mistake is designing for desktop screenshots first and users second.
A machine-readable page that humans hate is not a successful website.
WI-13 — Conversion and next action
Question
After a person finds the page, what job can the page complete?
Good:
- run the audit;
- start the tool;
- read the next technical guide;
- copy the spec;
- request a quote.
Bad:
14 competing CTAs and a chatbot asking whether you'd like a demo of the chatbot.
Conversion does not make you citable.
It makes citations valuable.
WI-14 — Knowledge graph and entity clarity
Question
Is it obvious what thing this page is about?
You want consistent:
- product name;
- company name;
- author identity;
- official URLs;
-
sameAs; - organization relationships;
- product naming;
- documentation naming.
If one page says:
- AuditMe;
- Audit ME;
- Audit-Me SEO Checker;
- AuditMe Website Intelligence;
…a machine has to decide whether these are one entity.
Make the answer boring.
One canonical name wins.
AuditMe itself exposes:
WI-15 — AI Search Readiness (GEO)
Question
Can a generative system retrieve and quote a bounded claim?
A strong page usually includes:
- answer in the first 100 words;
- descriptive headings;
- tables;
- examples;
- dates on changing facts;
- units on metrics;
- source links;
-
is / is notboundaries; - a real author;
- an obvious canonical URL.
Do not confuse this with:
- “add
llms.txt”; - “add FAQ schema”;
- “repeat keyword 30 times.”
Google's 2026 guidance explicitly emphasizes valuable, unique, non-commodity content and says there are no additional technical requirements or special schema types required for AI Overviews and AI Mode. Google: AI features and your website · Google: AI optimization guide
AuditMe:
WI-16 — Agent Readiness (ARO)
Question
Can an agent discover, parse and use the site without a human translating screenshots into instructions?
A strong implementation can provide some combination of:
- clean HTML;
- markdown representations;
-
llms.txt; - public API;
- documented API;
- MCP;
- predictable JSON;
- structured errors;
- clear authentication;
- rate-limit behavior;
- stable identifiers.
This is where the distinction between document retrieval and tool use becomes operational.
AuditMe:
Agent Safety Profile: readable is not the same as safe
A website can be beautifully machine-readable and still be hostile to an agent.
The agent-facing web introduces failure modes that classic SEO audits barely consider: prompt-injection text embedded in retrieved content, tools that mix read and write permissions, APIs that can be called in loops, expensive crawl endpoints, and credentials accidentally exposed through error messages.
WIS-2026 should therefore treat agent safety as a sibling of agent readiness, not as an optional security appendix.
The minimum ARO security contract
An agent-facing endpoint should make its boundaries obvious:
READ
inspect
search
audit
fetch
WRITE
create
update
delete
publish
Read operations should be safe to retry. Write operations should require explicit authorization and, where repeated execution could duplicate work, an idempotency mechanism.
A serious implementation should document:
- authentication and authorization;
- scopes/permissions;
- request and response size limits;
- timeout behavior;
- concurrency limits;
- rate-limit semantics and
Retry-Afterbehavior; - deterministic error codes;
- logging and audit trails;
- SSRF and URL-validation protections for user-supplied URLs;
- secret/credential handling;
- how untrusted page content is separated from tool instructions.
Rate limiting for agents
Do not invent a universal “10 requests per second” rule. That would become folklore immediately.
Instead, the standard should require implementations to protect finite resources and define limits deterministically. A production system may meter by IP, authenticated principal, API key, tool, tenant, cost budget, or concurrency class.
Expensive operations should not share the same budget as cheap reads:
GET /status cheap
GET /docs cheap
GET /audit?url= medium
POST /deep-crawl expensive
POST /render expensive
POST /publish write / high risk
Return a real machine-readable failure instead of a generic HTML “Oops!” page.
MCP-specific production reality
The MCP 2026-07-28 release introduced a stateless core, header-based routing, cacheable list results, authorization hardening, a formal extensions framework and updated SDKs. Those changes make ordinary web infrastructure — gateways, WAFs and rate limiters — more relevant to MCP deployments than ever. MCP 2026-07-28 release · Specification
That leads to a useful WIS rule:
If a website exposes an agent capability, its reliability contract is part of the website.
A tool that works only when nobody calls it is not agent-ready.
Open scoring: how to build a score that means something
A score without a formula is a mood.
WIS-2026 therefore publishes its assumptions.
6.1 Layer weights
The proposed public layer mix is:
| Layer | Weight |
|---|---|
| SEO | 35% |
| AEO | 15% |
| GEO | 25% |
| ARO | 25% |
Why?
Because:
- SEO remains the retrieval floor;
- AEO is the answer-extraction bridge;
- GEO measures source/citation value;
- ARO measures the move from reading to doing.
These weights are part of the proposal, not an industry fact.
A different implementation can use the taxonomy with different weights, provided it publishes them.
6.2 Dimension scores
Each dimension should be normalized to 0–100.
A simple layer model is:
layer_score =
sum(dimension_score × dimension_weight)
/ sum(applicable_dimension_weights)
The important word is applicable.
If a check cannot be evaluated, do not quietly treat “unknown” as “fail” or “pass.”
6.3 N/A is not zero
This is one of the easiest ways to corrupt an audit.
Suppose a page is a pure editorial article.
The absence of ecommerce inventory is not a failure.
Suppose a local shop has no public API.
That may be a lower ARO maturity, but it is not evidence of a broken local website.
So:
N/A ≠ 0
unknown ≠ fail
not applicable ≠ missing
A defensible score excludes genuinely non-applicable checks from the denominator.
6.4 Evidence provenance
Every numeric claim should state:
- What was measured.
- Unit.
- Method.
- Timestamp.
Example:
LCP: 15,310 ms — Lighthouse lab run — desktop — September 20, 2026.
Good.
Bad:
“Your page is very slow.”
Even if true, the second statement cannot be re-run.
6.5 Score bands
A useful example:
| Score | Interpretation | First move |
|---|---|---|
| 90–100 | Very strong floor | defend + improve evidence |
| 80–89 | Strong | fix the highest-recovery problems |
| 65–79 | Mixed | name the structural hole before publishing more content |
| 40–64 | Weak | repair technical/content foundations |
| 0–39 | Critical | restore ACCESS / CRAWL / INDEX first |
These are operating bands, not universal market tiers.
6.6 Recovery matters more than vanity score
A 78 can be more actionable than a 92 if the report tells you:
Current: 78
Recoverable: +14
Primary causes:
1. Canonical mismatch
2. Missing entity page
3. JS-only pricing content
That is a plan.
A report that says:
SEO score: 92
GEO score: 88
Congratulations!
…is decoration.
Internationalization: one entity, many language representations
Multilingual websites create a subtle GEO problem: translation is not the same thing as identity.
You want:
ONE ENTITY
├── English page
├── Spanish page
├── German page
└── Portuguese page
not:
FOUR PAGES THAT ACCIDENTALLY LOOK LIKE FOUR DIFFERENT COMPANIES
Google recommends using hreflang to connect localized versions and requires reciprocal references for the variants being declared. It supports HTML, HTTP headers and sitemap annotations. Google: localized versions
A WIS-compatible multilingual audit should therefore check:
- each locale has a stable canonical URL;
- each localized page identifies its own locale correctly;
-
hreflangalternatives form a valid reciprocal graph; -
x-defaultis used where a real fallback page exists; - visible product/entity names remain consistent across languages;
- structured data does not accidentally identify a translation as a different organization;
- language versions do not contradict pricing, availability, authorship or dates.
What about llms.txt?
Do not invent a standard requirement such as /llms.es.txt.
llms.txt is a proposal/convention defined by llmstxt.org, not a Google Search protocol with an official language-file hierarchy. A multilingual site can provide language-specific documentation paths or localized maps if that helps its own agent ecosystem, but WIS should label such patterns as implementation conventions, not universal requirements.
The correction problem: when the AI remembers the wrong thing
You cannot issue a universal command that rewrites the pretrained weights of every model that has ever seen your brand.
What you can control is the public truth surface that current retrieval systems can discover and verify.
Use explicit state transitions:
CURRENT
SUPERSEDED
DEPRECATED
CORRECTED
DISPUTED
UNKNOWN
For example, a pricing page can publish a machine-readable current offer while a changelog records:
2026-10-01
Old plan name retired.
New plan name: ...
Old price: ...
Current price: ...
Canonical source: ...
This does not magically purge stale model memory. It gives search/RAG systems, human researchers and agents a strong, dated correction trail.
A new metric: correction latency
WIS can measure:
Correction Latency = time between publishing an authoritative correction and observing the corrected claim in the target retrieval surface.
That is more honest than claiming you “updated the AI.”
For recurring AI visibility research, add two related metrics:
- Citation Persistence — how long a source remains cited across repeated observations;
- Entity Accuracy — the share of sampled responses that correctly identify the entity's current name, product, price, URL or other bounded facts.
Those are observations, not universal model properties. Report the engine, model/surface where known, locale, prompt, date and methodology.
The WIS-2026 scoring model, made implementable
The original standard proposes four layer weights:
| Layer | Weight |
|---|---|
| SEO | 35% |
| AEO | 15% |
| GEO | 25% |
| ARO | 25% |
The point is not that 35/15/25/25 is a universal truth.
It is that an open scoring method should publish the formula so another tool can reproduce the logic and challenge the weights.
The original WIS-2026 article explicitly frames these as proposed public weights rather than a claim about a search engine's algorithm.
18.1 Dimension score
Let:
d_i ∈ [0, 100]
be the score of dimension i.
Let:
w_i
be its normalized weight.
Then:
dimension_layer_score =
Σ(d_i × w_i)
for applicable dimensions.
But that formula alone is not enough.
18.2 N/A must not silently become zero
Suppose a site has no ecommerce catalog.
A product feed dimension should not automatically become:
0 = catastrophic failure
when the correct state is:
N/A = not applicable
A robust scoring implementation needs three states:
PASS
FAIL
N/A
N/A should be excluded from the denominator of the affected metric unless the methodology explicitly says otherwise.
This is one of the biggest places where audit scores become dishonest.
18.3 Evidence confidence
For high-stakes facts, add an evidence confidence field:
{
"score": 92,
"evidence_confidence": "high",
"evidence": [
{
"type": "html",
"timestamp": "2026-10-01T12:04:00Z"
}
]
}
Possible levels:
high
medium
low
unknown
This is not a search-engine ranking signal.
It is a reporting-quality signal.
18.4 Potential recovery
The question executives actually want answered is often:
“What can we realistically recover?”
So report:
current score
potential recovery
critical blockers
estimated effort
evidence confidence
Example:
Current: 74
Potential recovery from identified fixes: +14
Critical blocker: accidental noindex on 7 money pages
Highest leverage: indexability
This is more actionable than:
SEO health: 74/100
18.5 Score invariants
A serious implementation should test invariants such as:
same dimension id → same meaning
same evidence → same result
N/A → excluded as documented
score never < 0
score never > 100
layer weights sum to 100
overall score reproducible from stored inputs
If a score changes because an unrelated UI component changed, you don't have a stable audit engine.
4. The Machine-Readable Web: Robots, llms.txt, HTML & APIs
Robots.txt in 2026: search, training and user-triggered fetches are different jobs
This is the section most likely to prevent an unnecessary disaster.
There is no single category called “the AI bot.”
There are different jobs.
SEARCH / CITATION
≠
USER-TRIGGERED FETCH
≠
TRAINING
7.1 Googlebot
Googlebot crawls for Google Search.
Google documents that robots.txt can control crawling and that blocking Googlebot affects Google Search and related Search products. It also warns that blocking crawling is different from removing a URL from the index. Googlebot
Verify suspicious traffic by checking:
- user agent;
- source IP;
- reverse DNS.
Do not trust a string like Googlebot just because it appears in a header.
7.2 Google-Extended
This one causes endless confusion.
Google documents Google-Extended as a robots.txt product token, not a separate HTTP crawler user agent.
Google says it can be used to control whether content crawled from a site may be used for:
- training future Gemini model generations;
- grounding in Gemini Apps;
- Grounding with Google Search in Vertex AI.
Google also explicitly says Google-Extended:
- does not affect inclusion in Google Search;
- is not a Google Search ranking signal. Google common crawlers
So:
Google-Extended ≠ Googlebot
Google-Extended ≠ “the Google AI crawler”
7.3 OpenAI
OpenAI currently documents several user-agent roles.
OAI-SearchBot
Used for search.
OpenAI says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search results, though they can still appear as navigational links.
GPTBot
Used to crawl content that may be used for training OpenAI foundation models.
Disallowing GPTBot tells OpenAI not to use the site's content for that training purpose.
ChatGPT-User
Used for certain user-triggered actions.
OpenAI states that it is not used for automatic web crawling and, because these fetches are user initiated, robots.txt rules may not apply in the same way.
OpenAI also notes that robots changes for search may take roughly 24 hours to propagate through its systems. OpenAI: Overview of OpenAI Crawlers
That gives you an explicit control model:
OAI-SearchBot → search
GPTBot → training
ChatGPT-User → user action
Do not block all three just because “AI” feels scary.
7.4 Perplexity
Perplexity similarly separates jobs.
Its current documentation describes:
-
PerplexityBotfor surfacing and linking websites in Perplexity search; -
Perplexity-Userfor user-triggered fetches.
Perplexity says its search bot is not used to crawl content for AI foundation model training. It also recommends allowing the bot if you want to appear in Perplexity search results. User-triggered fetches generally ignore robots rules because the fetch was requested by a user. Perplexity crawler documentation
Again:
search access
and
training policy
are separate decisions.
7.5 A sane robots.txt example
This is an example policy, not a universal recommendation:
# Search engines
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /
# AI search / citation indexes
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Training policy example:
# keep search access while opting out of specific training crawls
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
# Private sections
User-agent: *
Disallow: /admin/
Disallow: /account/
Disallow: /api/internal/
Sitemap: https://www.example.com/sitemap.xml
The important part is not the exact file.
The important part is that the intent is explicit.
7.6 Do not use robots.txt to solve a noindex problem
If you want a URL not to appear in Google's index:
-
noindexis an indexing control.
If you block crawling with robots.txt, the crawler may never see the noindex.
Google explicitly distinguishes crawling, indexing and access controls. Googlebot
So:
robots.txt → can this crawler fetch?
noindex → should this URL be indexed?
auth → can users/crawlers access it at all?
Different job. Different tool.
7.7 WAF rules are part of SEO now
A perfect robots.txt can still be defeated by a WAF.
Common failure:
Googlebot → 200
Human browser → 200
OAI-SearchBot → 403 / challenge
PerplexityBot → 403 / challenge
The website owner then asks:
“Why doesn't AI cite me?”
The answer is not “add more GEO keywords.”
The answer is:
Your access policy is contradicting your visibility goal.
Perplexity's documentation explicitly recommends verifying both user agent and published IP ranges when configuring WAF rules. Perplexity bots
llms.txt: useful map, zero magic
llms.txt is one of the most misunderstood files in current AI web discussions.
The clean explanation is:
It is a curated map for LLM-oriented clients and agents.
It is not:
- a ranking factor;
- a replacement for sitemap;
- a substitute for good content;
- a way to override
noindex; - a guarantee that ChatGPT will cite you.
The public proposal is maintained at llmstxt.org.
Google's June 2026 documentation update explicitly clarified that llms.txt is not needed for Google Search and does not positively or negatively affect Search visibility or rankings. Google says it is fine to maintain for other services or systems that use it. Google Search documentation updates
That is the statement to use, not folklore.
8.1 What llms.txt is good at
A useful file can tell an agent:
- what the project is;
- which docs are canonical;
- which guides matter;
- where the API reference lives;
- where pricing lives;
- which pages explain the methodology.
This reduces discovery cost.
Think:
sitemap.xml
= complete inventory
llms.txt
= curated reading map
That difference is the whole point.
8.2 The v2 structure
The current llms.txt proposal defines a predictable Markdown shape.
A practical file contains:
# Product Name
> One-paragraph summary of what this project is and what matters.
Optional context and interpretation notes.
#### Docs
- [Quickstart](https://example.com/docs/quickstart): Start here.
- [API reference](https://example.com/docs/api): Full endpoint reference.
#### Guides
- [Architecture](https://example.com/docs/architecture): System overview.
- [Migration](https://example.com/docs/migration): Upgrade path.
The proposal's v2 guidance also recommends link relations for discovery of Markdown twins:
<link
rel="alternate"
type="text/markdown"
href="/docs/page.md"
>
<link
rel="describedby"
href="/docs/llms.txt"
>
Those same relationships can be expressed through an HTTP Link: header. llms.txt v2 specification
This is interesting because it moves beyond “one magical file” toward a broader idea:
A web page can explicitly advertise a machine-friendly representation of itself.
8.3 What not to put in llms.txt
Do not fill it with:
- every tag archive;
- every low-value page;
- tracking URLs;
- duplicate pages;
- broken links;
- marketing fluff;
- forty versions of the same guide.
A curated map should be curated.
A 200-line dump of your entire sitemap is not a map.
8.4 How to test your own llms.txt
Do not stop at:
curl -I https://example.com/llms.txt
Fetch the actual content:
curl -s https://example.com/llms.txt
Then check:
- H1 exists.
- Summary explains the project.
- URLs return 200.
- Important URLs are canonical.
- Link descriptions are useful.
- No junk dominates the file.
- An agent can identify where it should look first.
A simple human test:
Give a model only your
llms.txtand ask it to explain what the product does and where it should go for pricing, API, docs and support.
If the answer is wrong, the file failed its job.
8.5 AuditMe's own implementation
AuditMe maintains:
https://www.auditme.dev/llms.txt
The purpose is not “rank AuditMe.”
The purpose is to make the project's canonical resources easier for machine clients to discover.
That distinction is important enough to repeat:
A discovery file helps discovery. It does not manufacture authority.
Continuous Website Intelligence: put WIS in CI/CD
The modern failure is not “the team forgot SEO.” It is “a harmless frontend change silently broke machine access.”
A Next.js/Vercel deployment can regress:
canonical → missing
JSON-LD → removed
H1 → changed
robots → over-blocked
llms.txt → stale
sitemap → broken
raw HTML → empty
MCP endpoint → 500
pricing entity → inconsistent
A monthly audit catches this after the damage. CI catches it at the commit boundary.
The deployment pipeline
COMMIT
↓
UNIT TESTS
↓
BUILD
↓
STATIC WIS CHECKS
↓
DEPLOY PREVIEW
↓
HTTP SMOKE TEST
↓
HTML / ENTITY / JSON-LD TEST
↓
ROBOTS / SITEMAP / LLMS TEST
↓
MCP / API SMOKE TEST
↓
GOLDEN SNAPSHOT DIFF
↓
PASS → PRODUCTION
FAIL → BLOCK / REVIEW
A practical GitHub Actions command set might look like:
- run: npm test
- run: npm run build
- run: npm run wis:validate
- run: node scripts/check-robots.mjs
- run: node scripts/check-sitemap.mjs
- run: node scripts/check-llms.mjs
- run: node scripts/check-jsonld.mjs
- run: node scripts/check-hreflang.mjs
- run: node scripts/check-agent-surfaces.mjs
- run: npm run wis:smoke:preview
The exact scripts are implementation-specific. The principle is not.
Golden snapshots
Keep a small machine-readable “truth snapshot” for critical URLs:
{
"url": "https://example.com/",
"entity": "Example Corp",
"canonical": "https://example.com/",
"price": "$29",
"sameAs_count": 3,
"jsonld_types": ["Organization", "WebSite"]
}
A pull request that changes:
price: $29 → null
sameAs: 3 → 0
Organization JSON-LD: present → missing
should be visible before merge.
This is the website equivalent of type-checking a public API.
Severity policy
CRITICAL → deployment blocked
HIGH → review required
WARNING → deploy allowed, create issue
INFO → record only
The point is not to make every SEO nit a release blocker. The point is to make machine-critical regressions impossible to overlook.
Why this matters for WIS
The standard already ends with VERIFY → MONITOR. CI/CD closes the loop:
AUDIT
→ EVIDENCE
→ FIX
→ VERIFY
→ SNAPSHOT
→ DIFF
→ MONITOR
Website intelligence becomes an engineering discipline instead of a quarterly PDF ritual.
From HTML pages to machine-readable surfaces
The old web had a page.
The new web increasingly has representations.
For a high-value resource, consider:
Human HTML
│
├── semantic HTML
├── JSON-LD
├── canonical URL
├── Markdown twin
├── llms.txt relationship
└── optional API / MCP interface
This is not saying every site needs every layer.
It is saying different machines need different affordances.
9.1 HTML is still the canonical human document
Do not replace HTML with a text-only AI mirror.
Your primary page should still:
- work in browsers;
- be accessible;
- be indexable;
- explain the topic;
- show evidence.
The machine-friendly representation should reduce ambiguity, not become a secret second website.
9.2 Markdown twins
The llms.txt v2 proposal recommends explicit relations for Markdown versions.
Example:
https://example.com/docs/cache.html
https://example.com/docs/cache.html.md
This can be especially useful for documentation-heavy products.
The key is consistency.
The Markdown twin should not describe an entirely different product or outdated API.
9.3 Structured JSON
For functions and APIs, JSON becomes more valuable.
A useful response might look like:
{
"url": "https://example.com/",
"status": 200,
"score": 84,
"dimensions": {
"technical_seo": 91,
"content": 86,
"performance": 62,
"entity": 77
},
"evidence": [
{
"type": "http_status",
"value": 200,
"recorded_at": "2026-10-01T00:30:00Z"
}
]
}
The point is not the exact schema.
The point is that a machine can consume the result without parsing a screenshot.
9.4 Public API as an ARO signal
If your site's job is:
- price lookup;
- audit;
- availability;
- status;
- search;
- conversion;
- analytics retrieval;
…then an API is more useful than another “What is X?” article.
For an agent:
GET /api/v1/audit?url=https://example.com
can be a direct capability.
AuditMe's public interface:
5. Agent-Ready Websites: MCP, Capabilities & Security
MCP: when a website stops being a document and becomes a tool
The Model Context Protocol (MCP) is now a major part of agent infrastructure.
The current specification is 2026-07-28.
MCP defines a standardized way for AI applications to connect to external data sources and tools. Its protocol uses JSON-RPC messages, and servers can expose:
- resources;
- prompts;
- tools.
The 2026-07-28 release introduced a stateless protocol core, request/response patterns designed for ordinary HTTP infrastructure, cacheable list results, header-based routing and stronger authorization guidance. MCP specification · MCP 2026-07-28 release
That is important for websites because it makes the “agent interface” problem more concrete.
10.1 Document vs. capability
A human sees:
“Run a website audit.”
A document explains how to run one.
A tool exposes:
seo_audit(url)
The second can be executed.
That is the ARO distinction.
10.2 What the current MCP spec gives you
The 2026-07-28 specification defines:
resources → context/data
prompts → reusable instructions/workflows
tools → executable functions
The protocol uses self-contained requests and supports extensions such as Tasks and MCP Apps. MCP specification
For production systems, that also means taking security seriously.
MCP's own specification calls out:
- user consent;
- data privacy;
- tool safety;
- authorization;
- access controls.
It explicitly warns that tools may represent arbitrary code execution paths and that implementers need strong consent and authorization flows. MCP specification: Security and Trust & Safety
So “we have MCP” is not a complete ARO story.
“we have an authenticated, documented, safe MCP interface” is much closer.
10.3 ARO checklist for an MCP-enabled product
Discovery
- Is there documentation?
- Is the endpoint stable?
- Is there a clear description?
- Does the client know what the tools do?
Tool design
- Names are explicit.
- Inputs have machine-readable schemas.
- Outputs are structured.
- Errors are actionable.
- Side effects are obvious.
Security
- Public vs protected tools are deliberate.
- Auth is documented.
- Tokens are scoped.
- Sensitive actions require consent.
- Rate limits exist.
Reliability
- timeouts are bounded;
- retries are sensible;
- errors are structured;
- idempotency exists where appropriate;
- logs and traces exist.
10.4 AuditMe's own agent surface
AuditMe exposes:
- MCP documentation
- Claude setup
- Perplexity setup
- public endpoint at
https://www.auditme.dev/api/mcp - public REST documentation at https://www.auditme.dev/api-docs
The principle is:
Make the actual product job callable.
That is more durable than attaching “AI” to a dashboard label.
6. Entities, Structured Data, Provenance, i18n & Truth
Structured data, entities and provenance: make the page internally consistent
Structured data is often sold as if it were a secret handshake with Google, ChatGPT or some future agent.
It is not.
The useful mental model is simpler:
Your visible page is the source of truth. Structured data is the machine-readable index card describing that truth.
Google's structured-data documentation makes the same basic requirement clear: markup should represent the visible content on the page, and valid markup can make a page eligible for certain search features but does not guarantee that Google will show those features. See the Google structured data introduction and structured data policies.
That distinction matters even more in AI retrieval.
A model can survive a badly formatted <h2>. It is much harder to recover cleanly from a page that says one thing to a human, another thing in JSON-LD, and a third thing in its Open Graph metadata.
11.1 Build an entity graph, not a pile of schemas
For a typical SaaS website, the useful chain looks roughly like this:
Organization
│
├── Person (author / founder / expert)
│
├── WebSite
│ └── WebPage
│ ├── TechArticle
│ ├── BreadcrumbList
│ └── about → entities/topics
│
└── SoftwareApplication / Product
The exact vocabulary depends on the page. Do not copy this graph blindly.
A blog article may need TechArticle. A product page may need SoftwareApplication or another appropriate type. An organization page needs organization data. A breadcrumb trail may be useful where the page hierarchy actually has breadcrumbs.
The principle is:
one real-world entity → one stable identifier → consistent naming across the site.
11.2 Entity consistency checklist
Use the same canonical name in:
- the visible H1;
- title tag;
- Organization or Person markup;
- Open Graph title where appropriate;
- author bio;
- About page;
-
sameAslinks; - product navigation;
- documentation;
-
llms.txt; - API and MCP descriptions.
This does not mean every sentence should repeat the brand.
It means a crawler should not have to choose between:
AuditMe
Audit Me
AuditME
Audit-Me SEO Audit
AuditMe Website Intelligence
Pick one canonical public name and explain aliases only when they are genuinely used.
11.3 sameAs is not a trophy shelf
A common mistake is adding every social profile anyone has ever created.
sameAs is most useful when the URL actually identifies the same entity: an official GitHub organization or repository, an official social account, a founder profile, a recognized company page, or another durable identity surface.
Do not turn JSON-LD into a link dump.
A useful question is:
“Would a skeptical machine use this URL to disambiguate the entity?”
If the answer is no, leave it out.
11.4 about and mainEntity should describe the page you actually wrote
For a page about technical SEO audits:
{
"@type": "TechArticle",
"headline": "Technical SEO Audit: A 2026 Evidence-First Checklist",
"about": [
{
"@type": "Thing",
"name": "Technical SEO"
},
{
"@type": "Thing",
"name": "Website Audit"
}
],
"mainEntity": {
"@type": "Thing",
"name": "Technical SEO audit"
}
}
Do not declare a page's main entity as Product, Person or Organization just because that schema appears somewhere in the site template.
11.5 A modern article JSON-LD pattern
This is intentionally boring. Boring is good when machines need to understand the object.
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "TechArticle",
"@id": "https://example.com/blog/technical-seo-audit#article",
"headline": "Technical SEO Audit: A 2026 Evidence-First Checklist",
"description": "A practical checklist for crawling, indexing, rendering, links, structured data and performance.",
"datePublished": "2026-10-01",
"dateModified": "2026-10-01",
"inLanguage": "en",
"author": {
"@type": "Person",
"name": "Example Author",
"url": "https://example.com/about"
},
"publisher": {
"@type": "Organization",
"name": "Example",
"url": "https://example.com/"
},
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "https://example.com/blog/technical-seo-audit"
},
"about": [
{
"@type": "Thing",
"name": "Technical SEO"
}
],
"isPartOf": {
"@type": "Blog",
"@id": "https://example.com/blog"
}
},
{
"@type": "BreadcrumbList",
"@id": "https://example.com/blog/technical-seo-audit#breadcrumbs",
"itemListElement": [
{
"@type": "ListItem",
"position": 1,
"name": "Blog",
"item": "https://example.com/blog"
},
{
"@type": "ListItem",
"position": 2,
"name": "Technical SEO Audit",
"item": "https://example.com/blog/technical-seo-audit"
}
]
}
]
}
The important part is not the number of schema types.
It is internal agreement.
11.6 What changed with FAQ schema in 2026?
Do not publish this guide with the old assumption that adding FAQPage automatically creates a Google FAQ rich result.
Google removed documentation for the FAQ rich result and announced that the feature would no longer appear in Search starting May 7, 2026. The current Search Central updates explicitly say this. See Google Search documentation updates.
That does not mean FAQ content is useless.
Quite the opposite.
A good FAQ is still excellent human content and an excellent answer-extraction structure. What changed is the old rich-result incentive.
So:
FAQ content: useful
FAQ schema for a Google FAQ rich result: deprecated
FAQ schema as an excuse for weak content: useless
Write the questions because users and retrieval systems ask them. Mark up structured data only where the current documentation and the page content justify it.
The next web: from documents to capability graphs
Here is my prediction, and I am deliberately labeling it a prediction.
The web is not moving from “SEO” to “GEO” in a clean handoff. It is moving from documents toward a mixed system of documents, entities, datasets and callable capabilities.
A future website may expose four connected layers:
DOCUMENT GRAPH
pages, guides, docs, policies
ENTITY GRAPH
people, companies, products, places
EVIDENCE GRAPH
claims, sources, timestamps, measurements
CAPABILITY GRAPH
APIs, MCP tools, forms, transactions, live data
A human can read the document graph.
A search engine can retrieve from the document and entity graphs.
A generative system can cite the evidence graph.
An agent can traverse the capability graph and actually do something.
That suggests a new way to think about a homepage:
The homepage is not the entire website. It is the routing node for humans and machines.
It should answer:
WHAT IS THIS?
WHO IS BEHIND IT?
WHAT IS TRUE RIGHT NOW?
WHERE IS THE PRIMARY EVIDENCE?
WHAT CAN I READ?
WHAT CAN I CALL?
WHAT CAN I VERIFY?
The Website Intelligence Graph
I would extend WIS toward a future concept I call the Website Intelligence Graph (WIG):
Entity
│
├── canonical documents
├── localized documents
├── claims
├── evidence
├── external sources
├── current states
└── capabilities
│
├── API
├── MCP
├── search
├── transaction
└── live data
The interesting shift is that the unit of optimization is no longer the page.
It is the relationship between the page, the entity, the evidence and the capability.
That is where I think website intelligence is heading.
Three future metrics
If this model is right, the next wave of tooling will care about more than rankings and citations.
Evidence Reach — how many important user questions can be answered from authoritative evidence published by the entity?
Capability Success Rate — how often can an agent complete the intended task through the site's documented machine interface?
Truth Freshness — how quickly do the highest-value facts presented to retrieval systems converge on the site's current authoritative state?
These are proposed concepts, not established industry KPIs. The point of publishing them now is to make them falsifiable.
The rule that should survive every future model
Platforms will change.
Models will change.
Crawlers will change.
Protocols will change.
The durable rule is simpler:
Make the important thing easy to find, easy to understand, easy to verify, and — when appropriate — easy to call.
That is the future-facing core of WIS-2026.
7. Writing, Multimodal Search & Performance for Retrieval
Writing for retrieval: GEO content engineering without keyword soup
The internet has produced a weird genre of article that starts with:
“In today's rapidly evolving digital landscape…”
Then spends 1,300 words explaining why SEO matters before answering a one-sentence question.
That is bad writing for humans.
It is also bad material for retrieval systems.
A useful reference page should make important claims cheap to locate, quote and verify.
12.1 The first-100-words rule
For a core concept, answer the question immediately.
Bad:
Search has changed dramatically over the last few years. With the advent of artificial intelligence, many marketers are asking important questions about the future of organic visibility…
Better:
Generative Engine Optimization (GEO) is the practice of making content easier for AI answer systems to retrieve, understand, verify and cite. It is not a documented Google ranking factor named “GEO,” and it does not replace technical SEO.
The second version contains:
- the definition;
- the boundary;
- the misconception;
- the likely next question.
That is a retrieval-friendly opening.
12.2 Write sections that can survive extraction
Pretend a model will quote only one paragraph from your page.
Would that paragraph still make sense outside the page?
For every major concept, make a small answer unit:
Definition
→ direct answer
→ evidence
→ example
→ limitation
→ implementation
For example:
What is llms.txt?
llms.txt is a proposed convention for publishing a curated Markdown map of a website's most useful resources for language models and agents. It is not a replacement for robots.txt, and Google says it is not required for Google Search and has no positive or negative ranking impact. The current convention is documented at llmstxt.org and Google's clarification is in its Search documentation updates.
Notice what happened.
A reader can stop after one paragraph and still leave with a correct mental model.
12.3 Use tables where comparison is the actual job
Tables are unusually effective when the user is comparing objects.
Use them for:
- SEO vs AEO vs GEO vs ARO;
- Search crawler vs training crawler;
- lab vs field performance;
-
robots.txtvsnoindex; -
llms.txtvs sitemap; - HTML vs Markdown;
- REST API vs MCP;
- observation vs experiment;
- symptom vs likely cause vs first test.
Do not force a table into a narrative argument where a paragraph is clearer.
12.4 Dates belong next to changing facts
Compare:
“Google supports AI search.”
with:
“As of August 31, 2026, Google says the Search Console generative AI performance report is rolled out to websites worldwide.”
The second statement can be audited.
Google's current documentation says that the generative AI performance report provides data on a site's organic impressions from Search generative AI features, pages receiving those impressions, and their originating countries or devices. See the Search Console generative AI performance report.
For volatile facts, the date is part of the fact.
12.5 Units turn marketing into evidence
Never write:
“The site received massive visibility.”
Write:
“The site recorded 116,181 Google Search impressions and 9 clicks in the three-month first-party export used for this case study.”
Never write:
“AI citations increased a lot.”
Write:
“The separate AI-performance observation window recorded 1,737 citation events over 27 days.”
And always explain what the number means.
The source article that inspired this guide makes this provenance rule explicit: a numeric claim needs what was measured, its unit, the method, and a timestamp. That is one of the strongest parts of the original WIS-2026 framework.
12.6 “Is / is not” is one of the highest-value writing patterns
AI systems are very good at compressing related concepts into one fuzzy bucket.
Prevent that.
GEO is:
- content made easier to retrieve and cite in generative systems.
GEO is not:
- a documented Google ranking factor called “GEO”;
- a replacement for crawling and indexing;
- a guarantee of citation.
llms.txt is:
- a curated Markdown map.
llms.txt is not:
- robots.txt;
- a ranking signal documented by Google;
- a substitute for a sitemap.
Google-Extended is:
- a Google product token used in robots policy.
Google-Extended is not:
- a separate crawler user agent you should expect to see in logs.
This structure is useful because it blocks false equivalence.
12.7 Put examples beside abstract advice
Don't say:
“Use descriptive internal links.”
Show:
<a href="/blog/technical-seo-audit">
technical SEO audit checklist
</a>
rather than:
<a href="/blog/technical-seo-audit">
click here
</a>
Don't say:
“Expose an API.”
Show:
GET /api/v1/audit?url=https://example.com
Don't say:
“Document your bots.”
Show the relevant robots.txt policy and link the official documentation.
Examples compress ambiguity.
12.8 Write for query fan-out
Google says AI features can use a query fan-out technique: a complex user question can be decomposed into multiple related searches before a response is generated.
That suggests a practical content strategy.
A single strong reference page should contain independently useful sections for plausible subquestions.
Example topic:
“How to make a SaaS website visible in AI search in 2026”
Potential sub-intents:
- Does technical SEO still matter?
- How does Google AI Mode find sources?
- Should GPTBot be allowed?
- What is OAI-SearchBot?
- Does
llms.txtmatter? - How should JSON-LD represent a SaaS?
- How can you measure AI visibility?
- How can you make pages more citable?
- Does MCP help search visibility?
- What does an agent actually need from a site?
- How should a team monitor changes?
A 3,000-word page that answers every one badly is worse than a 6,000-word page that answers them precisely.
The goal is not word count.
The goal is retrievable coverage of the information need.
12.9 Semantic SEO is not synonym stuffing
Semantic coverage means covering concepts and relationships.
It does not mean:
SEO audit
SEO checker
SEO analysis
SEO analyzer
SEO inspection
SEO assessment
SEO evaluation
SEO review
SEO score checker
inside every paragraph.
Real semantic coverage would include:
crawlability
indexability
canonicalization
rendering
internal links
structured data
Core Web Vitals
entity clarity
author provenance
AI citations
robots policy
agent interfaces
The concepts create topical depth.
The synonyms do not.
The anti-AI-slop editorial standard
AI can help with research, extraction, clustering, first drafts and code.
That is not an argument for publishing generic AI text.
Google's 2026 documentation explicitly clarified that its spam policies apply to generative AI responses in Search, including large-scale generated content that does not add value. See Google's Search documentation updates.
A high-quality AI-assisted article should pass a stricter test:
Could a competent human have written this after actually doing the work, and can we show the work?
13.1 Never invent first-hand experience
Do not write:
“I tested 100 websites and discovered…”
unless you really did.
If you ran 17 tests, say 17.
If the result came from a vendor document, say that.
If a claim is based on a hypothesis, label it a hypothesis.
13.2 Never manufacture precision
There is no need to say:
“Adding FAQ schema increases citations by 2.8x.”
unless you have a reproducible experiment supporting it.
Better:
“We have not established a causal citation lift from FAQ schema. Google removed FAQ rich results from Search in 2026. We use FAQ sections primarily because they structure direct user questions.”
That sentence is less exciting.
It is much more defensible.
13.3 Separate four evidence classes
Use a visible evidence taxonomy:
| Evidence class | Example | How to write it |
|---|---|---|
| Official fact | Google documentation says X | “Google documents…” |
| First-party measurement | GSC export, server logs, audit output | “Our dataset recorded…” |
| Experiment | controlled before/after test | “In this test…” |
| Interpretation | strategic conclusion | “A practical interpretation is…” |
That small distinction dramatically improves credibility.
13.4 Show the failure cases
A page that only documents success is a marketing page.
A reference guide should contain misses:
- what did not move;
- what looked like a ranking problem but was actually indexability;
- what looked like a GEO issue but was entity ambiguity;
- what looked like a schema problem but was content quality;
- what looked like a bot problem but was a WAF challenge.
The original AuditMe article does this well with its failure matrix. Keep expanding that pattern.
13.5 Publish the methodology
For every original test, record:
experiment:
id: GEO-2026-001
date_start: 2026-09-01
date_end: 2026-09-14
pages:
- https://example.com/page-a
engines:
- ChatGPT
- Perplexity
- Gemini
prompts:
count: 20
change:
variable: "definition block"
comparison:
control: "original page"
treatment: "answer-first rewrite"
outcome:
citation_rate_before: 0.20
citation_rate_after: 0.35
caveats:
- "small sample"
- "engine behavior can change"
That is more useful than a screenshot.
13.6 Keep a changelog
Technical reference pages decay.
When a search engine changes a crawler policy, a reporting interface, a schema feature, or an API contract, a dated changelog tells both humans and machines that the document is alive.
Example:
##### Changelog
- 2026-10-01 — Updated Google AI Search measurement guidance; removed deprecated FAQ rich-result language.
- 2026-09-24 — Added multimodal Search Console reporting.
- 2026-09-01 — Added Bing AI Performance notes.
- 2026-07-28 — Updated MCP references to the 2026-07-28 specification.
A current page should not silently retain a 2024 recommendation.
Multimodal search: the image is becoming a query
Text-only SEO is an incomplete model.
On September 24, 2026, Google announced web multimodal Search performance reporting in Search Console for searches involving Lens, Circle to Search on Android, image uploads to Google Search, and Chrome's “Search this image” flow. The rollout began globally that day. See Google Search Central: web multimodal Search performance reporting.
That changes the practical question for image-heavy sites:
Can the visual object be understood, associated with the right page, and retrieved when the image itself becomes the user's query?
14.1 Image SEO basics still matter
For meaningful images:
<img
src="/images/technical-seo-crawl-map.webp"
alt="Technical SEO crawl map showing discovery, indexing, and canonicalization"
width="1600"
height="900"
loading="lazy"
/>
The exact implementation depends on where the image appears.
The principles are stable:
- useful alternative text;
- stable image URL;
- correct dimensions;
- sensible compression;
- meaningful surrounding text;
- no accidental layout shift;
- an image that actually illustrates the claim.
14.2 Alt text is not a keyword field
Bad:
alt="seo seo audit seo checker auditme best seo tool 2026"
Better:
alt="Website audit dashboard showing crawl, content, schema and performance findings"
Describe the image.
Do not describe the keywords you wish somebody would search.
14.3 The page around the image provides context
A product image without surrounding product data is ambiguous.
A chart without a caption is ambiguous.
A screenshot without an explanation is ambiguous.
For high-value visuals, consider:
**Figure 1. Website retrieval pipeline.**
The diagram shows the transition from HTTP access through crawl,
discovery, indexing, retrieval, citation and conversion.
That text can become useful context for humans, search systems and multimodal retrieval.
14.4 Consistency across text, image and data
Bing's AI Performance guidance also points publishers toward consistency across formats: text, images and other media should describe the same product, entity and concept.
That is common sense, but teams violate it constantly.
Example:
- H1: “AuditMe Website Intelligence”
- hero screenshot: old product name
- JSON-LD: unrelated
Product - image alt: “SEO report”
- OG title: “Free SEO Checker”
-
llms.txt: “AI SEO Audit”
The system has to guess what the object is.
Reduce the guesswork.
Performance is not GEO — but a slow site can still destroy retrieval
One of the laziest claims in SEO is:
“Core Web Vitals are a direct GEO ranking factor.”
There is no good basis for stating that as a general documented fact.
A more accurate model is:
Performance
↓
fetch reliability
↓
render reliability
↓
human usability
↓
ability to consume content
Performance belongs in the website intelligence model because speed and stability are important properties of a usable, retrievable website.
Google's current Web Vitals guidance uses these “good” p75 targets:
| Metric | Good target |
|---|---|
| LCP | ≤ 2.5 s |
| INP | ≤ 200 ms |
| CLS | ≤ 0.1 |
See web.dev: Web Vitals.
15.1 Field data and lab data are different measurements
This distinction should be printed in every serious report.
Lab data is a controlled synthetic measurement.
Field data reflects real users and real environments.
Do not put:
LCP: 15.3s
in a report without saying whether that is lab or field.
The original WIS-2026 article uses exactly this discipline: its AuditMe case study labels the 15.31-second figure as lab/PSI rather than pretending it is CrUX field data. That provenance-first treatment should stay.
15.2 Raw HTML vs rendered HTML
Modern frameworks make it possible for a page to look excellent after JavaScript executes while shipping very little useful HTML initially.
That creates a useful audit question:
What can a fetcher understand before it runs your entire application?
Check:
curl -sL https://example.com/page | sed -n '1,240p'
Then inspect:
-
<title>; - canonical;
- robots directives;
- H1;
- first paragraphs;
- important links;
- JSON-LD;
- primary content.
For JavaScript-heavy sites, compare that with the rendered DOM in a browser.
The discrepancy is not automatically a bug.
But a large discrepancy should be intentional.
15.3 The LCP fix should be specific
“Improve performance” is not a fix.
A useful audit says:
Problem:
Hero image blocks LCP.
Evidence:
1.4 MB AVIF/JPEG asset is the largest render-critical resource.
Action:
- serve a correctly sized responsive image
- preload only if justified
- eliminate unnecessary third-party blockers
- verify mobile field/Lab results after deployment
Verification:
rerun PageSpeed + inspect field data after enough traffic accumulates
That is a fix.
8. Testing & Measurement: From Manual Audit to CI/CD
Technical cookbook: audit the machine-readable web yourself
The fastest way to become better at SEO/GEO is to stop relying on dashboards for every question.
A surprising amount can be checked from the command line.
16.1 Check HTTP status and redirects
curl -I -L https://example.com/
Look for:
- status code;
- redirect chain;
- HTTPS;
- cache headers;
- server behavior.
You want to understand whether the “real” URL is obvious.
16.2 Fetch raw robots.txt
curl -s https://example.com/robots.txt
Then separately inspect the key tokens:
Googlebot
OAI-SearchBot
GPTBot
PerplexityBot
Google-Extended
Don't assume one User-agent: * block tells the whole story.
16.3 Check the sitemap
curl -s https://example.com/sitemap.xml
Then ask:
- Does it exist?
- Is the XML valid?
- Are URLs canonical?
- Are there obvious 404s?
- Does it contain junk archives?
- Does lastmod reflect meaningful changes?
A sitemap is a discovery aid, not a guarantee of indexing.
16.4 Check llms.txt
curl -s https://example.com/llms.txt
A practical first pass:
[ ] returns 200
[ ] Markdown is readable
[ ] one clear H1
[ ] concise summary
[ ] canonical resource links
[ ] no dead URLs
[ ] no giant marketing essay
[ ] no private URLs
The current llms.txt convention is documented at llmstxt.org. Google explicitly says Google Search does not require llms.txt and it does not have a positive or negative ranking effect. Keep it for systems that may use it, not as a fake Google lever. See Google's Search documentation updates.
16.5 Extract title, H1 and canonical quickly
On a machine with common Unix tools:
curl -sL https://example.com/page \
| grep -Ei '<title>|<h1|rel=["'\'']canonical'
For a serious audit, use a real HTML parser rather than regex.
The point of the shell command is reconnaissance.
16.6 Extract JSON-LD
curl -sL https://example.com/page \
| grep -o '<script[^>]*type=["'\'']application/ld\+json["'\''][^>]*>.*</script>'
Again: for production tooling, parse the DOM.
The audit question is:
Does the machine description agree with the human page?
16.7 Verify bot identity correctly
User agents can be spoofed.
Google's crawler documentation recommends verification using source IP and reverse DNS rather than trusting a UA string alone. See Google: verifying Googlebot.
For OpenAI and Perplexity, use their published crawler/IP documentation when configuring WAF rules:
- OpenAI bot documentation
- OpenAI published IP lists
- OpenAI SearchBot IPs
- Perplexity bot documentation
Do not hard-code an old blog post's IP list into a firewall and forget it exists.
16.8 Compare access policy with business intent
Create a simple matrix:
| Surface | Allow? | Why? |
|---|---|---|
| Google Search | Yes/No | Organic search |
| ChatGPT search | Yes/No | Search/citations |
| Perplexity search | Yes/No | Search/citations |
| AI training | Yes/No | Training policy |
| Private admin | No | Security |
| Private API | No | Security |
| Public documentation | Yes | Discovery/use |
This is far more useful than copying somebody else's robots.txt.
Measuring AI visibility in 2026: stop using one screenshot as analytics
The biggest change in AI visibility is not that we suddenly know everything.
It is that several systems now expose parts of the picture.
17.1 Google Search Console
As of August 31, 2026, Google says its generative AI performance insights are rolled out to all websites worldwide.
The report can show:
- impressions from Search generative AI features;
- pages receiving those impressions;
- countries;
- devices.
See Google Search Console: Generative AI performance report.
The important phrase is impressions from generative AI features.
An impression is not a click.
A citation is not a conversion.
A view in an AI response is not proof that the user read your page.
Measure each thing separately.
17.2 Bing Webmaster Tools
Bing now has an AI Performance report covering supported Copilot, Bing AI summaries and selected partner experiences.
It exposes citation-oriented metrics such as:
- total citations;
- cited pages;
- average cited pages;
- grounding queries;
- page-level citation activity;
- trends.
It also has preview concepts such as Intents, Topics, Citation Share and Compare. See Bing AI Performance and Bing's June 2026 announcement.
Bing explicitly warns that citation activity is not the same thing as ranking, authority or importance.
That warning should be copied into your own dashboard.
17.3 Prompt panels
For systems where first-party reporting is limited, use a frozen prompt panel.
A useful panel might contain:
prompt_suite:
informational:
- "what is [category]?"
- "best tools for [job]?"
- "how do I solve [problem]?"
commercial:
- "[category] software comparison"
- "[brand] alternatives"
- "best [category] tool for [use case]"
entity:
- "what is [brand]?"
- "who makes [product]?"
technical:
- "how does [feature] work?"
Run the same prompts on a consistent schedule.
Record:
date
engine
prompt
answer
cited domains
cited URLs
brand mentioned?
brand position/context?
claim correct?
link present?
A screenshot is useful for evidence.
It is not sufficient as a time series.
17.4 Citation rate
One simple internal metric:
citation rate =
prompt runs where target domain was cited
/
total eligible prompt runs
Example:
17 cited runs / 50 total runs = 34%
Do not call 34% a “ranking.”
It is a measurement of your defined prompt panel.
17.5 Supported-claim rate
A stronger quality metric:
supported-claim rate =
AI claims about your company that can be verified
/
all sampled claims about your company
This catches a failure mode that citation counts miss.
You can be cited and still be described incorrectly.
17.6 Citation quality
A citation should be evaluated on:
- correct entity — is this actually your page?
- correct claim — does the page support what the answer says?
- current evidence — is the source stale?
- useful destination — can the reader act on it?
- context — did the system cite you for the primary fact or a minor side note?
That makes “we were cited” much less vanity-driven.
17.7 Agent success rate
For ARO, use an operational metric:
agent success rate =
successful task completions
/
attempted tasks
Example tasks:
- fetch a pricing page;
- retrieve API documentation;
- execute a public read-only query;
- find a specific technical article;
- call a documented audit endpoint;
- interpret structured JSON.
A website can have excellent citations and terrible agent usability.
That is not contradictory.
9. Diagnostics, Failure Modes & Industry Playbooks
A 25-failure diagnostic matrix
When an audit says “something is wrong,” start with the failure state rather than the feature you were hoping to sell.
| Symptom | Likely cause | First verification |
|---|---|---|
| Indexed page gets no AI mentions | Weak answer shape / low relevance | Inspect first 100 words + query fit |
| AI cites old pricing | Stale documentation | Compare cited URL content to current offer |
| Search impressions high, clicks tiny | Query mismatch / low rank / SERP behavior | Inspect query-level GSC data |
llms.txt exists, nothing changes |
No guaranteed Search effect | Check actual citations, not file presence |
| ChatGPT does not cite | Wrong bot / weak content / retrieval mismatch | Check OAI-SearchBot policy + prompt panel |
| Perplexity does not cite | Bot blocked / content mismatch | Check PerplexityBot + WAF |
| Googlebot works, AI fetch fails | WAF or fetch policy | Review access logs |
| Schema valid, entity wrong | Markup contradicts visible page | Compare JSON-LD to H1/byline/price |
| FAQ rich result disappeared | Feature was deprecated | Review Google current docs |
| Mobile LCP awful | Large hero / third-party JS | PSI waterfall |
| Raw HTML nearly empty | Client-only rendering |
curl vs rendered DOM |
| Sitemap huge and dirty | Auto-generated archive URLs | Crawl + sample status codes |
| Many pages indexed, few cited | Commodity content | Compare uniqueness and evidence |
| Many citations, zero conversions | Answer completes task | Add tool/data/next-step value |
| Score high, users unhappy | Audit covers surface quality only | Check UX and conversion |
| Score unstable after code changes | Non-deterministic scoring source | Snapshot audit JSON + tests |
| 0 score for impossible feature | N/A mishandled | Inspect applicability logic |
| “AI score” generated by LLM | No deterministic evidence | Trace field provenance |
| LCP called “field data” | Lab/field confusion | Verify measurement source |
| Bot user agent spoofed | UA-only access rule | Verify IP |
| WAF challenge HTML returned 200 | Soft block | Inspect body, not status only |
noindex accidentally added |
Template or deployment | Inspect raw HTML/meta |
| Canonical points elsewhere | Duplicate URL architecture | Compare canonical vs intended URL |
| Outbound sources all dead | Link rot | Status crawl |
| Content updated without changelog | Stale evidence perception | Add timestamp and change record |
This table is intentionally diagnostic, not magical.
One symptom may have multiple causes.
The audit should narrow the hypothesis.
Industry playbooks
The fundamentals stay the same, but the implementation priorities change.
20.1 SaaS companies
A SaaS site should be understandable as an entity and usable as a product surface.
Core pages:
Home
Product
Use cases
Pricing
Docs
API
Security
Status
About
Methodology / trust
Comparisons
High-value WIS dimensions:
- entity clarity;
- technical SEO;
- content quality;
- structured data;
- AI search readiness;
- agent readiness;
- API availability;
- conversion.
A surprisingly common SaaS mistake is building a beautiful homepage with no durable documentation.
AI can recommend a product.
Then it follows the citation to:
“Request a demo”
and cannot understand pricing, limits, API behavior or integration details.
Make the evidence public where business policy allows.
Useful AuditMe resources:
20.2 Ecommerce
Ecommerce has a richer structured-data and product-data problem.
The core stack is:
product page
+ visible price/availability
+ Product structured data
+ consistent product data
+ high-quality images
+ Merchant Center where relevant
+ accurate variants
+ reviews/content where policy allows
Do not put a product in JSON-LD that the human page does not show.
Do not let price drift between:
page
schema
feed
checkout
Multimodal discovery also matters more for visually distinctive products. Google's 2026 multimodal Search reporting makes image-based discovery measurable in Search Console.
20.3 Documentation and developer tools
Developer sites are unusually suited to ARO.
They can expose:
- stable docs;
- Markdown;
- OpenAPI;
- API examples;
- machine-readable schemas;
- SDKs;
- MCP;
- changelogs;
- status;
- versioned docs.
A good API page might have:
What it does
Endpoint
Authentication
Parameters
Example request
Example response
Errors
Rate limits
Version
Last updated
A model does not need 2,000 words of marketing copy.
It needs the contract.
20.4 Local businesses
Local search has an entity problem:
same business
same name
same address
same phone
same hours
same category
Keep those consistent.
For local businesses, the conversion layer matters enormously because the user may want:
- directions;
- hours;
- appointment;
- phone;
- reservation;
- menu;
- availability.
A beautiful local SEO article is less useful than an accurate page containing the thing the customer actually needs.
20.5 Publishers and media
For publishers:
- author identity;
- dates;
- original reporting;
- editorial policy;
- corrections;
- source links;
- article update history;
- canonicalization;
- image attribution.
The biggest GEO advantage is often not “AI optimization.”
It is being the original source.
Google's 2026 interface changes around preferred sources and provenance reinforce the importance of identifiable original content, but do not turn “preferred source” into a guaranteed traffic mechanism.
20.6 Agencies
Agency reporting should stop producing:
SEO score: 87
without showing:
What was measured
What failed
Why it matters
Evidence
Recommended action
Expected recovery
Verification
AuditMe's workflow is deliberately built around this:
audit
→ evidence
→ priority
→ fix
→ verify
See the AuditMe methodology and monthly SEO reporting guide.
10. Implementation: 30 Days, 90 Days, Templates & JSON
The 30-day Website Intelligence sprint
You do not need a quarter to discover whether the basics are broken.
Days 1–3: establish the baseline
Run:
[ ] technical crawl
[ ] indexability check
[ ] robots review
[ ] sitemap review
[ ] canonical review
[ ] raw HTML inspection
[ ] structured-data validation
[ ] CWV baseline
[ ] entity consistency check
[ ] AI search prompt baseline
Save the artifacts.
Do not overwrite them.
Days 4–7: repair ACCESS → INDEX
Fix only blockers:
HTTPS
robots
noindex
canonicals
redirects
sitemap
orphan pages
rendering
Do not publish 40 blog posts while the canonical points to a 404.
Week 2: make the important pages extractable
For your 10 highest-value URLs:
[ ] answer in first 100 words
[ ] clear H1
[ ] useful H2 questions
[ ] evidence tables
[ ] dates and units
[ ] author / entity
[ ] sources
[ ] visible examples
[ ] concise FAQ
[ ] one clear next action
Week 3: build the AI surface
Decide your policy:
Search:
allow / block?
Training:
allow / block?
User-triggered:
allow / block?
llms.txt:
publish / maintain?
Markdown:
publish / maintain?
API:
public / authenticated?
MCP:
needed / unnecessary?
No folklore.
Base the decision on business intent.
Week 4: measure
Repeat the prompt panel.
Open:
- Google Search Console generative AI report;
- Bing AI Performance;
- server logs;
- citation records;
- conversion analytics.
Then compare the period.
Do not claim causation from correlation.
Write:
“After the content change, citation rate rose from X to Y in our panel.”
not:
“This change increased ChatGPT ranking by 38%.”
The second statement implies a measurement you may not actually possess.
The 90-day operating model
The long-term goal is not one perfect audit.
It is a system that notices deterioration.
Month 1 — Foundation
crawl
index
render
performance
schema
entity
content
robots
measurement
Month 2 — Evidence
Publish:
- one original dataset;
- one original framework;
- one transparent experiment;
- one deep technical guide;
- one comparison page;
- one implementation reference.
Do not publish six rewrites of the same generic “What is GEO?” article.
Month 3 — Operations
Automate:
weekly:
- HTTP checks
- robots diff
- sitemap diff
- selected URL crawl
- CWV trend
- prompt panel
- citation log
- changelog
Then treat the outputs as an engineering backlog.
Copy-paste templates
23.1 Evidence block
> **Evidence**
>
> **Observed:** 1,247 organic impressions
> **Period:** 2026-09-01 → 2026-09-30
> **Source:** Google Search Console
> **Metric:** Search impressions
> **Method:** first-party export
> **Caveat:** impressions are not clicks or conversions
23.2 Experiment block
##### Experiment: answer-first rewrite
**Hypothesis:** A definition in the first 100 words increases retrieval usefulness.
**Control:** original page.
**Treatment:** rewritten opening with direct definition, boundary and source.
**Measurement:** citation appearance in a frozen prompt panel.
**Period:** 2026-09-15 → 2026-09-29.
**Result:** [record the actual result].
**Limitations:** small sample; engine behavior changes; no causal claim beyond the test design.
23.3 Methodology block
##### Methodology
This page combines three evidence types:
1. First-party measurements from AuditMe.
2. Current official documentation from search engines and protocols.
3. Explicitly labeled interpretation and recommendations.
Dates are included for changing platform behavior. Lab and field performance data are not mixed.
23.4 Citation log
date,engine,prompt,target_domain,target_url,cited,correct,link_present,notes
2026-10-01,ChatGPT,"what is website intelligence?",auditme.dev,https://www.auditme.dev/blog/website-intelligence-standard-2026,true,true,true,"direct citation"
2026-10-01,Perplexity,"how to audit AI readiness",auditme.dev,https://www.auditme.dev/blog/website-intelligence-standard-2026,false,false,false,"competitor cited"
23.5 llms.txt template
# Example Product
> Example Product is a [one sentence factual description].
#### Docs
- [API documentation](https://example.com/docs/api): endpoints and authentication.
- [Developer guide](https://example.com/docs): implementation reference.
#### Guides
- [Technical SEO guide](https://example.com/blog/technical-seo): crawl and indexing guidance.
- [Methodology](https://example.com/methodology): scoring and evidence rules.
See llmstxt.org for the current convention.
23.6 HTML links for a Markdown representation
The current llms.txt convention also documents link relations for alternative Markdown representations. A pattern can look like:
<link
rel="alternate"
type="text/markdown"
href="/docs/guide.md"
title="Guide in Markdown"
/>
<link
rel="describedby"
type="text/markdown"
href="/docs/guide.md"
/>
Treat this as a machine-readable representation strategy, not a Google ranking trick.
23.7 A public API response
{
"url": "https://example.com/",
"generated_at": "2026-10-01T12:00:00Z",
"standard": "WIS-2026",
"scores": {
"overall": 82,
"seo": 84,
"aeo": 77,
"geo": 80,
"aro": 86
},
"dimensions": [],
"provenance": {
"method": "url_audit"
}
}
Notice what is missing:
“AI likes this site: 92”
A report should expose measurable objects, not vibes.
JSON reference: WebsiteIntelligenceReport
A practical report object should make provenance first-class.
{
"$schema": "https://www.auditme.dev/specs/website-intelligence-report-2026.json",
"standard": "WIS-2026",
"standard_url": "https://www.auditme.dev/blog/website-intelligence-standard-2026",
"generated_at": "2026-10-01T12:00:00Z",
"input": {
"url": "https://example.com/",
"final_url": "https://example.com/",
"http_status": 200
},
"scores": {
"overall": 82,
"band": "strong",
"potential_recovery": 13
},
"layers": {
"seo": {
"score": 84,
"unit": "rank_click"
},
"aeo": {
"score": 77,
"unit": "extract"
},
"geo": {
"score": 80,
"unit": "citation_event"
},
"aro": {
"score": 86,
"unit": "retrieve_or_tool_call"
}
},
"engine": {
"name": "ExampleAuditEngine",
"checks_run": 159,
"dimensions": 16
},
"evidence_policy": {
"numeric_claims_require": [
"what",
"unit",
"method",
"timestamp"
]
},
"surfaces": {
"robots_txt": true,
"sitemap": true,
"llms_txt": true,
"markdown": false,
"api": true,
"mcp": true
},
"dimensions": [
{
"id": "WI-15",
"name": "AI Search Readiness",
"score": 88,
"status": "pass",
"applicability": "applicable",
"evidence": [
{
"type": "first_100_words",
"recorded_at": "2026-10-01T12:00:00Z"
}
],
"fixes": []
}
],
"provenance": {
"method": "url_audit",
"lab_and_field_separated": true
}
}
For a production schema, publish an actual JSON Schema file and version it.
A URL that says “latest” forever is not an API contract.
Use explicit versions:
/specs/website-intelligence-report-2026.json
or:
/specs/wis-2026/report.schema.json
Then document breaking changes.
11. How to Become a Citable Source — and Where the Web Goes Next
How to make a page citable without trying to game a model
There is no reliable “ChatGPT SEO trick.”
There are reliable information-design practices.
25.1 Become the source
The strongest content assets are things you actually know because you did the work:
- original datasets;
- benchmarks;
- code;
- experiments;
- methodology;
- product telemetry;
- transparent case studies;
- primary interviews;
- direct testing.
AuditMe's own Website Intelligence Standard is more defensible as a source because it defines a named framework, methodology and report shape rather than saying “here are 17 GEO hacks.”
25.2 Name your framework
A named framework creates an identifiable object.
Website Intelligence Standard 2026
WIS-2026
Then define it consistently.
But don't confuse naming something with establishing authority.
A name is useful when the underlying artifact is real and maintained.
25.3 Publish counterarguments
A strong reference page should answer:
“Where could this model be wrong?”
For WIS-2026:
- Search engines can change retrieval systems.
- Citation panels sample the AI web.
- AI answers vary by query, time, location and system state.
- MCP does not automatically create search traffic.
-
llms.txthas no documented Google ranking benefit. - a high audit score does not guarantee market share.
- a citation does not guarantee a click.
These caveats make the methodology more believable.
25.4 Keep claims granular
Bad:
“This architecture makes your site AI-ready.”
Better:
“This architecture exposes the primary page content in HTML and publishes a Markdown representation, reducing the amount of client-side execution required to read the content.”
The second statement can be checked.
What AuditMe should not claim
This section is deliberately uncomfortable.
AuditMe should not claim that it:
- knows Google's secret ranking formula;
- can guarantee ChatGPT citations;
- can guarantee Google AI Overview inclusion;
- replaces Ahrefs' historical backlink index;
- replaces GA4 attribution;
- proves causality from a single prompt;
- turns
llms.txtinto a ranking factor; - turns MCP into an SEO factor;
- knows what every AI system will answer tomorrow.
Instead, say exactly what is measured.
That is how a product becomes a source rather than another SEO score generator.
The difference between “visible” and “useful”
This is the heart of the whole framework.
Imagine four websites.
Website A
Ranks well.
But the answer is hidden behind:
JS bundle
→ modal
→ login
→ interactive widget
Search may tolerate some of that.
An agent may not.
Website B
Is cited constantly.
But every citation points to a generic homepage with no evidence.
It gets mentioned.
It does not get trusted.
Website C
Has perfect technical SEO.
But every article is a rewritten summary of public sources.
The site is crawlable.
It is not original.
Website D
Has fewer pages.
But each important page contains:
definition
evidence
example
sources
date
entity
API/docs
clear next action
That is the object WIS-2026 is trying to measure.
Not “AI score.”
Machine usefulness.
A practical homepage audit in 15 minutes
Pick the homepage.
Open:
https://example.com/
Then answer these questions.
Identity
What is this company?
What is the product?
Who is it for?
Can you answer in one sentence without scrolling?
Evidence
What does the product actually do?
How do you know?
Where is the documentation?
Retrieval
Does the first HTML response contain the core explanation?
Entity
Is the company/product name consistent?
Does schema agree?
Search
Is it indexable?
Canonical?
Linked internally?
AI
Can a model quote a precise paragraph?
Are claims dated?
Are sources linked?
Agents
Can software discover the docs?
Can it call the important public function?
Are errors machine-readable?
Conversion
What should the user do next?
If you cannot answer those questions quickly, a “perfect SEO score” is not the thing you need.
The machine-readable homepage test
Try this thought experiment:
A model gets only the raw HTML,
robots.txt, sitemap,llms.txt, JSON-LD and public docs.
Can it reconstruct:
{
"entity": "...",
"product": "...",
"audience": "...",
"core_capability": "...",
"primary_urls": [],
"pricing": "...",
"documentation": "...",
"api": "...",
"authoritative_sources": []
}
If the answer is no, your problem may not be “GEO.”
It may simply be poor information architecture.
The retrieval-first content template
Use this structure for a flagship reference article.
# Primary topic
> **Direct answer:** one concise paragraph.
#### TL;DR
- definition
- key facts
- practical result
#### What is [topic]?
Definition + boundary.
#### Why it matters in 2026
Current facts with dates and sources.
#### How it works
A simple system diagram.
#### Evidence
Original data + methodology.
#### What it is not
Explicit boundaries.
#### Step-by-step implementation
Actionable instructions.
#### Examples
Real, reproducible examples.
#### Failure modes
Symptom → cause → verification → fix.
#### Measurement
What to record and what not to claim.
#### FAQ
Short direct answers.
#### Sources
Primary documentation first.
#### Changelog
Dated updates.
#### Related AuditMe tools
Natural next actions.
This is a better SEO/GEO strategy than writing 5,000 words first and deciding what the article means at paragraph 37.
The “citation-worthiness” test
Before publishing any important section, ask:
Could a journalist, engineer, SEO, developer or AI system cite this paragraph without adding a missing qualifier?
Test it.
Claim
llms.txtimproves AI visibility.
Problem
Which visibility?
Which system?
Observed or documented?
What does “improves” mean?
Rewrite
llms.txtis a curated Markdown map that may help systems that consume the convention discover a site's preferred resources. Google says the file is not needed for Google Search and does not positively or negatively affect Search visibility or rankings.
Now the claim has boundaries.
That's GEO writing.
Not keyword density.
Final framework: the Website Intelligence equation
A useful conceptual equation is not:
SEO score + GEO score + AI score = 100
It is:
ACCESS
×
CRAWL
×
DISCOVER
×
INDEX
×
RETRIEVE
×
UNDERSTAND
×
CITE
×
USE
×
CONVERT
×
VERIFY
Not mathematically literally.
Operationally.
If any critical stage fails, the downstream stage may never get a chance.
A page cannot be cited if it cannot be retrieved.
A page cannot be retrieved reliably if it cannot be discovered.
A page cannot be understood cleanly if its visible content and machine representation disagree.
A citation does not create value if the page cannot finish the user's job.
And a result you cannot reproduce is not a reliable result.
That is the purpose of WIS-2026.
Closing: the website is becoming an interface to knowledge
SEO did not disappear.
The web object got larger.
A serious website in 2026 is simultaneously:
a search document
a source document
an entity
a data surface
a visual object
a human interface
and sometimes a tool
Google still has to retrieve pages.
AI systems still need evidence.
Users still need answers.
Agents increasingly need capabilities.
The mistake is pretending those are the same operation.
They are not.
That is why the Website Intelligence Standard 2026 separates:
SEO → rank / click
AEO → extract
GEO → citation
ARO → retrieve / tool call
The dimensions can overlap.
The measurement should not.
The evidence should not.
The claims should not.
And the score should never be allowed to become more important than the artifact behind the score.
Search visibility is useful. Citation visibility is useful. Agent readiness is useful.
But the end state is simpler:
Make the website the easiest truthful source to understand, verify and use.
That is a strategy that survives the next crawler name.
12. AuditMe Reference, Checklists, Sources, FAQ & Release
AuditMe as a reference implementation
The easiest way to understand WIS-2026 is to separate the standard from the runtime.
The standard
Defines:
taxonomy
dimensions
layer questions
weights
evidence policy
JSON contract
The runtime
Does:
fetch
parse
crawl
score
produce evidence
prioritize
recheck
AuditMe is the runtime for this standard.
Start here
- AuditMe home
- Free SEO Analyzer
- Website SEO Checker
- Audit Methodology
- AI Visibility
- Research
- API docs
- MCP docs
- MCP for Claude
- MCP for Perplexity
llms.txt- GitHub: AISeoAudit
Relevant AuditMe research
- Your Website Was Seen 116,181 Times and Clicked 9 Times
- How to Measure AI Search Visibility in 2026
- What Actually Makes ChatGPT, Claude & Perplexity Cite Your Website
- Generative Engine Optimization (GEO) 2026
- How AI Systems Read the Web in 2026
- AI Readiness:
llms.txt, robots.txt and semantic HTML - Complete SEO Audit Guide 2026
- SEO Audit Checklist 2026
- Technical SEO fundamentals
- JavaScript SEO and rendering
- Core Web Vitals checklist
- JSON-LD guide
- Internal links strategy
- WordPress SEO audit
- SEO audit API guide
- Best SEO checker tools 2026
The source article already treats these pages as a cluster around the WIS-2026 hub. Keep that architecture, but make each child page genuinely deeper rather than creating dozens of near-duplicate keyword pages.
SEO + GEO + ARO technical checklist
Crawlability
[ ] HTTPS
[ ] stable 200 response
[ ] redirect chain sane
[ ] robots policy deliberate
[ ] sitemap available
[ ] crawlable internal links
[ ] no accidental noindex
[ ] canonical valid
Content
[ ] clear H1
[ ] answer early
[ ] unique substance
[ ] descriptive headings
[ ] examples
[ ] dates
[ ] units
[ ] sources
[ ] author
[ ] update history
AEO
[ ] definitions are direct
[ ] questions are explicit
[ ] lists are extractable
[ ] tables are structured
[ ] concise answers
[ ] no contradictory sections
GEO
[ ] entity is stable
[ ] claims are bounded
[ ] original evidence exists
[ ] source links are primary
[ ] citation panel exists
[ ] AI references are measured, not guessed
ARO
[ ] raw HTML is useful
[ ] Markdown representation considered
[ ] llms.txt maintained if useful
[ ] bots policy explicit
[ ] API documented
[ ] MCP considered where function calls matter
[ ] structured errors
[ ] authentication documented
Performance
[ ] field vs lab labeled
[ ] LCP measured
[ ] INP measured
[ ] CLS measured
[ ] mobile tested
[ ] critical resources understood
Integrity
[ ] no invented numbers
[ ] no unsupported “ranking factor” claims
[ ] no stale bot policy
[ ] no fake freshness
[ ] no schema contradiction
[ ] no SEO score theatre
Official documentation map
Use primary sources before blog posts.
Google Search
- AI features and your website
- AI optimization guide
- Google Search documentation updates
- Generative AI performance report
- Search generative AI control
- Web multimodal Search performance reporting
- Googlebot
- Overview of Google crawlers
- Verifying Googlebot
- Structured data introduction
- Structured data policies
- Core Web Vitals
Robots standard
OpenAI
Perplexity
llms.txt
MCP
Bing
- AI Performance in Bing Webmaster Tools
- AI Performance announcement
- Intents, Topics, Citation Share and Compare
A source hierarchy for SEO and GEO research
When deciding which source to trust, use a hierarchy.
Tier 1 — primary specification
Examples:
- RFC;
- protocol specification;
- official search documentation;
- official API documentation.
Use these for “how the system works.”
Tier 2 — first-party measurements
Examples:
- Google Search Console;
- Bing Webmaster Tools;
- server logs;
- CrUX;
- Lighthouse / PSI;
- your own experiment logs.
Use these for “what happened on this site / dataset.”
Tier 3 — original research
Examples:
- reproducible tests;
- named datasets;
- transparent methodology;
- controlled experiments.
Use these for “what we observed.”
Tier 4 — reputable secondary analysis
Useful for context, examples and interpretation.
Tier 5 — social posts and anecdotes
Useful for finding questions.
Weak as proof.
The fact that a claim has been repeated 400 times on LinkedIn does not make it true.
How to reference sources without killing the reader
A giant bibliography at the bottom is useful.
Inline links are still better.
Bad:
AI search is changing SEO.[37][41][52]
Better:
Google says AI Overviews and AI Mode continue to rely on its Search infrastructure and that pages need to be indexed and eligible to display snippets; see Google's AI features documentation.
The link tells the reader exactly where the evidence lives.
For a long technical article, use:
claim → inline source
then maintain a consolidated source map at the bottom.
That gives both humans and machines a clear path.
SEO content architecture for AuditMe
This article should act as a hub.
A useful internal-link graph is:
WIS-2026
│
┌──────────────────┼──────────────────┐
│ │ │
Technical GEO/AEO ARO
│ │ │
crawl/index citations llms.txt
CWV entities APIs
links content MCP
schema measurement agent docs
│ │ │
└────────────── AuditMe ──────────────┘
│
Free Analyzer
The hub should answer the broad conceptual question.
Child pages should answer narrower implementation questions.
This prevents the classic problem where 60 articles all attempt to rank for:
“AI SEO”
while none becomes the canonical resource for a specific problem.
Recommended keyword and entity coverage for this page
Do not stuff these terms into every paragraph.
Cover them naturally where they belong:
website intelligence
SEO audit
technical SEO
SEO checker
SEO analyzer
SEO score
AEO
Answer Engine Optimization
GEO
Generative Engine Optimization
AI search
AI visibility
AI citations
AI Overviews
AI Mode
Bing AI Performance
ChatGPT search
Perplexity
Googlebot
OAI-SearchBot
GPTBot
PerplexityBot
Google-Extended
robots.txt
llms.txt
MCP
Model Context Protocol
agent readiness
agentic web
structured data
JSON-LD
schema.org
entity SEO
knowledge graph
Core Web Vitals
LCP
INP
CLS
technical audit
content quality
E-E-A-T
provenance
crawlability
indexability
rendering
internal links
canonical tags
sitemap
multimodal search
AI agents
API
machine-readable website
The surrounding concepts matter more than exact-match repetition.
Meta title, description and social snippet
Recommended page title:
The Website Intelligence Standard 2026: SEO, GEO, AEO & Agent Readiness
Alternative social headline:
Google Ranked You. AI Cited You. Agents Skipped You.
Recommended meta description:
A 2026 evidence-first framework for SEO, AEO, GEO and Agent Readiness: crawling, AI citations, robots.txt, llms.txt, structured data, MCP, Core Web Vitals, measurement and practical fixes.
Keep the canonical:
https://www.auditme.dev/blog/website-intelligence-standard-2026
For DEV.to syndication, point the canonical back to AuditMe.
The source article already establishes this canonical strategy and explicitly describes syndication as distribution rather than a second original.
FAQ for Search, AI systems and humans
What is the Website Intelligence Standard 2026?
WIS-2026 is a proposed open, evidence-first framework for evaluating a website across four layers: SEO, AEO, GEO and ARO (Agent-Readiness Optimization). It defines 16 dimensions, scoring rules, evidence requirements and a machine-readable report model.
Is SEO dead in 2026?
No. Google's current guidance says the same foundational SEO practices remain relevant to AI features in Search, and pages need to be indexed and eligible to appear in supporting links. SEO is still the retrieval floor.
What is GEO?
Generative Engine Optimization is the practice of making content easier for generative answer systems to retrieve, understand and cite. “GEO” is not a documented Google ranking factor with that name.
What is AEO?
Answer Engine Optimization focuses on making information easy to extract as a direct answer, such as a concise definition, list, table or supporting passage.
What is ARO?
Agent-Readiness Optimization focuses on making a website usable by software agents: discoverable documentation, machine-readable content, clear robots policy, stable APIs and tool interfaces where appropriate.
Does llms.txt improve Google rankings?
Google says llms.txt is not needed for Google Search and has no positive or negative effect on Search visibility or rankings. It remains a convention that some other systems may choose to consume.
Is llms.txt the same as robots.txt?
No. robots.txt is a crawler access policy. llms.txt is a curated content map intended for language models and agents.
Does llms.txt guarantee ChatGPT citations?
No. There is no general guarantee. A citation depends on retrieval, relevance, source selection and system behavior.
Is Google-Extended a crawler?
No. Google describes Google-Extended as a product token used in robots.txt controls for Gemini-related training and grounding. It is not a separate crawler user agent.
Should I block GPTBot?
That is a publisher policy choice. OpenAI distinguishes GPTBot, which is associated with training, from OAI-SearchBot, which is associated with search. You can make separate choices.
Should I allow OAI-SearchBot?
If ChatGPT search visibility is part of your publishing objective, blocking the search crawler makes the content less available to that search path. Confirm the current OpenAI documentation before changing policy.
What is ChatGPT-User?
OpenAI documents ChatGPT-User as a user-triggered fetch mechanism. Treat user-triggered retrieval separately from search indexing and training.
What is PerplexityBot?
Perplexity documents PerplexityBot as its search crawler. It is distinct from Perplexity-User, which handles user-triggered requests.
Does JSON-LD make ChatGPT cite a page?
There is no documented rule that JSON-LD directly causes ChatGPT citations. Structured data can reduce ambiguity about entities and page structure when it accurately represents the visible content.
Does Google have special AI Overview schema?
Google says there is no special schema required specifically for AI features. Use supported structured data accurately where appropriate.
Does FAQ schema still create a Google FAQ rich result?
No. Google removed the FAQ rich-result feature from Search in 2026.
How do I measure AI Overviews?
Use the Google Search Console generative AI performance report for Google's generative AI features, then use a separate frozen prompt panel and other first-party or platform-specific telemetry for systems that do not expose equivalent reporting.
Does Bing measure AI citations?
Yes. Bing Webmaster Tools has an AI Performance report that exposes citation-oriented metrics for supported Microsoft/Copilot and partner experiences.
Is a citation the same as a click?
No. A citation means content was referenced or shown as a source. It does not prove a click, session or conversion.
Can a site be cited but receive no traffic?
Yes. AI answers can satisfy a user's question without requiring them to open the source.
Do Core Web Vitals affect AI citations?
Do not treat that as an established direct citation factor. Performance matters for usability and reliable fetching, but citation systems are not publicly reducible to a Core Web Vitals formula.
What are the 2026 Core Web Vitals thresholds?
At the p75 “good” level, LCP is ≤2.5 seconds, INP ≤200 milliseconds and CLS ≤0.1.
Do I need MCP?
Only when agents need to perform a function that is better exposed as a tool than retrieved as a document. A documentation article does not need MCP simply because MCP is fashionable.
Does MCP improve SEO?
Not as a documented Google ranking factor. MCP is a tool interface. Its direct value is enabling machine-to-machine interaction where the product actually has a callable function.
Does a public API make a website more agent-ready?
It can. The strongest case is when the API exposes the core task the product performs, with clear documentation, authentication rules, predictable errors and stable responses.
Is raw HTML important in an agent-ready website?
Usually yes, especially for public information pages. The easier the primary content is to retrieve without executing a large client application, the fewer assumptions a simple fetcher needs to make.
What should an SEO audit report contain in 2026?
At minimum: what was measured, evidence, severity, recommended fix, priority, method, timestamp and verification state. A score without those fields is incomplete.
What is WIS-2026 not?
It is not Google's ranking formula, not a guarantee of AI citations, not a backlink database, not an analytics replacement and not a prediction of market share.
Can I implement WIS-2026 without AuditMe?
Yes. The framework is designed as a public method. The runtime is optional.
What is AuditMe?
AuditMe is a website intelligence and SEO/GEO audit platform that turns a URL into evidence and a prioritized fix list across technical SEO, content, performance, structured data, accessibility, security, entity clarity, AI search readiness and agent readiness.
The one-page operational checklist
Before publishing any flagship SEO/GEO page, verify:
IDENTITY
[ ] one canonical entity name
[ ] author identified
[ ] organization identified
[ ] about / mainEntity makes sense
SEARCH
[ ] URL returns 200
[ ] indexable as intended
[ ] canonical correct
[ ] internal links exist
[ ] sitemap includes it
[ ] robots policy deliberate
CONTENT
[ ] answer in first 100 words
[ ] clear H1
[ ] question-led H2s
[ ] original evidence
[ ] dates and units
[ ] examples
[ ] limitations
[ ] primary sources
[ ] FAQ written for humans
STRUCTURED DATA
[ ] valid JSON-LD
[ ] matches visible page
[ ] no obsolete rich-result assumptions
AI SEARCH
[ ] AI-retrievable claims
[ ] entity clarity
[ ] citation-worthy evidence
[ ] prompt panel baseline
AGENTS
[ ] raw HTML readable
[ ] docs discoverable
[ ] llms.txt considered
[ ] API documented if relevant
[ ] MCP considered if relevant
[ ] robots separates business policies
PERFORMANCE
[ ] field vs lab labeled
[ ] LCP
[ ] INP
[ ] CLS
[ ] mobile
DISTRIBUTION
[ ] canonical set
[ ] OG title
[ ] OG description
[ ] OG image
[ ] syndication canonical back to source
INTEGRITY
[ ] no fake numbers
[ ] no unsupported ranking claims
[ ] no fake freshness
[ ] changelog present
Sources and further reading
This guide intentionally prefers first-party documentation.
Search and AI
- Google: AI features and your website
- Google: AI optimization guide
- Google Search documentation updates
- Google Search Console: Generative AI performance
- Google: Web multimodal Search reporting
- Google: Preferred Sources
Technical web
- Googlebot
- Google crawler overview
- Googlebot verification
- Robots Exclusion Protocol — RFC 9309
- Web Vitals
- Google structured data
- Schema.org
AI crawlers
Machine-readable web
AI visibility measurement
- Bing AI Performance
- Bing AI Performance announcement
- Bing Intents, Topics, Citation Share and Compare
AuditMe
- AuditMe
- Free SEO Analyzer
- Website SEO Checker
- Audit methodology
- AI Visibility
- Research
- API docs
- MCP docs
llms.txt- GitHub repository
AuditMe research cluster
- Website Visibility Gap
- AI Search Visibility measurement
- 47 citation tests
- GEO Visibility Guide
- AI Web Intelligence
- AI Readiness Guide
- Complete SEO Audit Guide 2026
- Technical SEO fundamentals
- JavaScript SEO
- Core Web Vitals
- JSON-LD
- Internal links
- SEO audit API
- Best SEO checker tools 2026
Changelog
-
2026-10-01 — Expanded the WIS-2026 reference into a comprehensive SEO × AEO × GEO × ARO field guide. Added current Google Search Console generative AI reporting, multimodal Search reporting, Bing AI Performance, updated crawler distinctions, MCP 2026-07-28, structured-data guidance,
llms.txtlimitations, measurement methodology, implementation templates, industry playbooks, and failure diagnostics. - 2026-10-01 — Removed any assumption that FAQ schema creates Google FAQ rich results. Google ended that feature on May 7, 2026.
-
2026-10-01 — Clarified that
llms.txtis a convention, not a documented Google Search ranking mechanism. - 2026-10-01 — Kept SEO, GEO and ARO claims explicitly separated from measured first-party evidence.
How to cite this page
Tymchenko, E. (2026, October 1). The Website Intelligence Standard 2026: SEO, AEO, GEO and Agent-Readiness as one measurable stack. AuditMe. https://www.auditme.dev/blog/website-intelligence-standard-2026
Canonical URL:
https://www.auditme.dev/blog/website-intelligence-standard-2026
JSON-LD starter for the published article
Use a version like this on the actual page and keep every field synchronized with what is visible.
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "TechArticle",
"@id": "https://www.auditme.dev/blog/website-intelligence-standard-2026#article",
"headline": "The Website Intelligence Standard 2026: SEO, AEO, GEO and Agent-Readiness as one measurable stack",
"alternativeHeadline": "Google Ranked You. AI Cited You. Agents Skipped You.",
"description": "A 2026 evidence-first framework for measuring website visibility and machine usability across SEO, AEO, GEO and agent readiness.",
"datePublished": "2026-10-01",
"dateModified": "2026-10-01",
"inLanguage": "en",
"author": {
"@type": "Person",
"name": "Eduard Tymchenko",
"url": "https://www.auditme.dev/about"
},
"publisher": {
"@type": "Organization",
"name": "AuditMe",
"url": "https://www.auditme.dev/",
"sameAs": [
"https://github.com/Edo911/AISeoAudit",
"https://dev.to/edo911"
]
},
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "https://www.auditme.dev/blog/website-intelligence-standard-2026"
},
"about": [
{
"@type": "Thing",
"name": "Search Engine Optimization"
},
{
"@type": "Thing",
"name": "Generative Engine Optimization"
},
{
"@type": "Thing",
"name": "Agent-Readiness Optimization"
},
{
"@type": "Thing",
"name": "Website Intelligence"
}
],
"isPartOf": {
"@type": "Blog",
"@id": "https://www.auditme.dev/blog"
}
},
{
"@type": "BreadcrumbList",
"@id": "https://www.auditme.dev/blog/website-intelligence-standard-2026#breadcrumbs",
"itemListElement": [
{
"@type": "ListItem",
"position": 1,
"name": "Blog",
"item": "https://www.auditme.dev/blog"
},
{
"@type": "ListItem",
"position": 2,
"name": "Website Intelligence Standard 2026",
"item": "https://www.auditme.dev/blog/website-intelligence-standard-2026"
}
]
}
]
}
Do not add FAQPage markup just to chase a Google rich result. FAQ content can remain as normal visible content; Google says the FAQ rich-result feature is no longer shown in Search.
Machine citation contract
Canonical topic: Website Intelligence Standard 2026 (WIS-2026)
Publisher: AuditMe
Primary author: Eduard Tymchenko
Status: proposed open framework
Core model: SEO + AEO + GEO + ARO
Core dimensions: 16
Evidence rule: claim → method → artifact → source → timestamp → scope
Primary workflow: audit → evidence → diagnosis → priority → fix → verify → monitor
Important non-claims: WIS-2026 is not a Google ranking formula;llms.txtis not presented as a documented Google ranking factor; no technique in this document guarantees AI citation.
Canonical article: https://www.auditme.dev/blog/website-intelligence-standard-2026
When citing WIS-2026, prefer the canonical URL above and preserve the version/date of the edition you relied upon.
The final test
Before you call this page “done,” run three tests.
The human test
Can an experienced SEO, developer or founder find an actionable answer in under two minutes?
The citation test
Can a model quote the important claims without inventing missing context?
The engineer test
Can someone reproduce the major measurements from the documented method?
If the answer to all three is yes, the page is doing something worthwhile.
If the answer to any is no, add evidence — not adjectives.
This document is a proposed open standard and field guide, not a search-engine ranking formula. Platform behavior changes. Re-check linked official documentation before treating a dated crawler, API, search feature or protocol behavior as current.
Primary implementation: AuditMe — URL in, evidence out.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.