I scored my own AI-citation tool against 4 strangers' pages. It failed on mine.
Two days ago I published an AI Citation Readiness Checker. Eight conditions a page should meet before a generative search engine might cite it, each traced to a named public study. It runs in the browser, nothing is uplo
Two days ago I published an AI Citation Readiness Checker. Eight conditions a page should meet before a generative search engine might cite it, each traced to a named public study. It runs in the browser, nothing is uploaded, no signup.
Then I wrote the second post about it and realised a problem I had not thought about.
Every time I had tested it, I had tested it on my own page.
A tool validated only on the author's own page cannot tell the difference between "my rubric is correct" and "I wrote the rubric to match my own page." Those look identical from inside. So I ported the eight checks to Node, thresholds unchanged, and ran them against four pages that have nothing to do with me.
The measurement
Same eight checks, same thresholds, same day (2026-10-11), on the raw server-returned HTML. No JavaScript executed, which is what an AI retriever sees.
| Page | Kind | Score | Failing |
|---|---|---|---|
| the-seo-autopilot.com indie SaaS benchmarks | external | 8/8 | — |
| dev.to "Developer Passive Income Reality Check" | external | 8/8 | — |
| blog.csdn.net distribution analysis | external | 7/8 | 1 |
| my own geoprobe page | mine | 7/8 | 1 |
| siwan.io indie hacker revenue report | external | 6/8 | 1 |
Here is the per-check detail, because the totals alone hide the interesting part:
| Page | Body text chars | List shape | Quantified claims | Internal links | Canonical | Date | AI crawlers |
|---|---|---|---|---|---|---|---|
| the-seo-autopilot | 19,293 | 93% (110 li / 8 p) | 147 | 95 | absolute | yes | allowed |
| dev.to longform | 18,408 | 40% (39 li / 58 p) | 29 | 30 | absolute | yes | allowed |
| CSDN analysis | 12,772 | 32% (39 li / 83 p) | 49 | 0 | absolute | yes | allowed |
| my geoprobe | 1,975 | 33% (2 li / 4 p) | 6 | 2 | absolute | yes | allowed |
| siwan.io | 9,406 | 0% (0 li / 31 p) | 8 | 17 | absolute | yes | not fetched |
All five passed the same four checks: server-rendered body, JSON-LD, absolute canonical, visible date. siwan.io's robots.txt did not come back this run, so I marked it unverified rather than counting it as "allowed".
What this actually told me
Prose essays lose to data tables, on the same rubric. siwan.io's revenue report is a good post. It has real numbers, it is exactly the topic I care about, and it scored the lowest of the five. Its problem was structural: 31 paragraphs, 0 list items. Nothing on that page can be quoted as a block. The two 8/8 pages are both dense data pages, 93% and 40% list-shaped.
The practical version: restructuring prose into tables is not cosmetic. It changes what a retriever can lift verbatim.
My own page failed, and it was the check I would never have run manually. The failure was "is this part of a topic cluster" — my checker page had 2 internal links, threshold is 3. That is a completely mundane omission. I would not have caught it by eye, ever, because a page with two links looks completely normal to the person who wrote it.
The fix was a uniform nav bar across every tool page. That took all seven of my pages from 6-7/8 to 8/8.
The check I could not fix was the one that mattered. My geoprobe page has 1,975 characters of body text against 9,000-19,000 for the external pages. I could have added tables and lists to push the ratio up. I did not, because that would be gaming a number I wrote myself. The tool measures readiness. It does not know the page is worth citing.
What this does not tell you
I have no citation-panel data for any AI engine. I cannot prove that any of these five pages, including mine, is actually being cited by anything. Readiness is a precondition, not an outcome.
Every check here can be gamed without creating any value. You can hit 93% list-shape with an empty table. The rubric tells you whether a page is structurally quotable. Whether it is worth quoting is a separate question this tool does not answer and cannot answer.
If you take one thing from this: a self-built tool needs at least one run against something you did not make. Mine took two posts and one wasted day to notice it had never been done.
The raw measurement, including every check message, is here as JSON: https://first-dollar-reality.app.workbuddy.host/bench/benchmark.json
The tool itself, if you want to run it on your own page: https://first-dollar-reality.app.workbuddy.host/geoprobe/index.html
One caveat on the benchmark page above: it is live, but the seven-page nav fix described here is local until I confirm the redeploy. I am not going to report a URL as fixed before curl says 200.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.