154 of 529 top homepages have no canonical tag. 15 point to a URL that redirects straight back.
I fetched the homepages of the Tranco top 1,000 once each and read two tags that tell crawlers which URL a page is: <link rel="canonical"> and the hreflang alternates. 529 homepages came back as real pages. 154 of the 52
I fetched the homepages of the Tranco top 1,000 once each and read two tags that tell crawlers which URL a page is: <link rel="canonical"> and the hreflang alternates. 529 homepages came back as real pages. 154 of the 529 have no canonical tag at all. Of the 375 that do, 317 point at the URL they were served from and 58 point somewhere else. I then fetched those 58 targets: 20 redirect, 15 of them straight back to the page that named them, one is a 404 and one is a 403.
The hreflang side is in better shape than its reputation. 197 homepages declare language versions. I fetched up to three alternates for each and checked whether they link back. 479 of the 522 alternates I fetched do; 15 of the rest didn't answer 200.
Below: the limits first, the numbers, the mistakes that repeat, and a short checklist.
Limits first (measured 2026-10-06/07)
- One fetch per domain, from Japan. Homepages on 2026-10-06 between 21:30 and 21:54 UTC, then the canonical targets and alternates between 21:55 and 22:03 UTC, one request at a time, 10 second timeout, redirects followed, no JavaScript. Several sites serve a Japanese page to a Japanese IP, and their canonical changes with it. From another country you'd see other URLs for some of these.
-
A bot user agent. I identified myself (
glovrex-measure/1.0). 96 domains answered 403 to that, which is fewer pages than my earlier runs with a browser user agent. Those sites are unknown here, not "no canonical". - List and set: Tranco list 56WKN, top 1,000. Many entries are CDN, API or tracking domains with no homepage: 314 gave no HTTP answer (DNS, TLS, timeout), 148 answered with something other than 200, 9 answered 200 with no HTML head or an "Access Denied" page. n = 529 below.
-
Homepage only,
<head>only. I read canonical and hreflang from the HTML head and from the HTTPLinkheader. hreflang in an XML sitemap isn't covered, so a site can be fine there and still show up as "no hreflang" here. -
Return links are sampled. At most three alternates per site, skipping
x-defaultand the page itself. - I don't know how AI crawlers use these tags. Google and Bing document canonical and hreflang. OpenAI, Anthropic and Perplexity don't say whether their crawlers read them. Answer engines that build on a search index inherit whatever the index decided.
The canonical numbers
n = 529 homepages.
| Homepages | |
|---|---|
No canonical (head or Link header) |
154 |
| Canonical = the URL served | 317 |
| Canonical points elsewhere | 58 |
| More than one canonical tag | 4 |
| Two canonicals that disagree | 1 |
Relative canonical (/, /us/) |
2 |
og:url present |
291 |
og:url disagrees with canonical |
19 |
The 58 that point elsewhere, by what differs:
| Difference | Homepages |
|---|---|
| Only the query string | 18 |
| A different host | 15 |
| A different path | 14 |
| Only a trailing slash | 8 |
Only www.
|
3 |
The query-string group is mostly fine. bing.com, ya.ru and youtu.be redirect with a tracking or redirect parameter and the canonical drops it, which is what a canonical is for.
What the 58 targets answered
| Canonical target | |
|---|---|
| 200, and declares itself canonical | 32 |
| Redirects to another URL | 20 |
| No answer | 2 |
| 404 | 1 |
| 403 | 1 |
| 200, no canonical of its own | 1 |
| 200, canonical points on again | 1 |
The 20 redirects are the interesting group. 15 of them redirect to exactly the URL that declared them, 3 more come back to it with a fresh query parameter (ya.ru twice, salesforce.com), and 2 go somewhere new (amp.dev, hihonorcloud.com). The loops look like this:
-
Trailing slash, both ways. huawei.com serves
/en/and names/enas canonical, and/enredirects to/en/. agora.io does the same. ubisoft.com does it in the other direction. -
www.both ways. go.com is served atwww.go.comwith canonicalgo.com, which redirects towww.go.com. bugsnag.com does the same. -
Mobile and edition hosts. vk.com sends my user agent to
m.vk.com, whose canonical isvk.com, which redirects back tom.vk.com. cnn.com lands onedition.cnn.comwith canonicalwww.cnn.com. -
A path the server won't serve. epa.gov names
/home, which redirects to/.
A canonical that redirects isn't fatal. Google calls rel=canonical and redirects both signals and picks one. But here the two signals point in opposite directions, so the site has handed the choice to the crawler.
The two hard failures are a stanford.edu homepage whose canonical is https://www.stanford.edu/home, which answers 404 (rechecked with curl, same result), and synology.com, whose canonical is https://www.synology.com// with a double slash, which answers 403.
The canonical depends on where you are
Six of the "different host" and "different path" cases, from four companies, point at a regional page when seen from Japan:
- facebook.com (and fb.com, fbcdn.net): canonical
https://ja-jp.facebook.com/ - android.com: canonical
https://www.android.com/intl/ja_jp/ - forbes.com: canonical
https://www.forbes.com/home_asia/ - intel.com: served from
www.intel.co.jp, canonical onwww.intel.com
This can be intentional. It also means a crawler in the US and a crawler in Japan are told different things about the same URL, and a homepage can quietly be declared a copy of a regional page. If you serve by location, check the canonical from more than one country.
og:url versus canonical
19 of 291 disagree. Most are small (instagram.com has canonical www.instagram.com and og:url instagram.com), but two are real: roblox.com's homepage says og:url is /CreateAccount, and gandi.net's canonical is /en-US while og:url is /en. Link previews and some crawlers read og:url, so it's worth keeping it in sync.
The hreflang numbers
n = 197 homepages with hreflang (37% of 529). Median 13 alternates, maximum 270.
| Homepages | |
|---|---|
Includes x-default
|
137 |
| Lists itself among its alternates | 183 |
| Doesn't list itself | 14 |
| hreflang on a page whose canonical points elsewhere | 19 |
| A code that isn't a valid language or region | 7 |
A 3-letter language code (fil, ceb, skr) |
5 |
A withdrawn code (iw, in) |
4 |
| Same code, two different URLs | 4 |
| Relative or non-HTTP URL | 2 |
The 7 invalid ones: weebly.com and amp.dev use underscores (en_GB, pt_BR); viber.com uses ua and gr, which are country codes, not languages; aliyun.com uses tc; hp.com writes country first (py-es, us-en) plus internal names like lamerica_nsc_carib-en; ezviz uses region eu; shein.com uses hr-eur and zh-tw-sg. Google's localized versions guide asks for an ISO 639-1 language with an optional ISO 3166-1 region. I count fil or iw separately because they are valid language tags, but fil has no ISO 639-1 code and iw was replaced by he in 1989.
"hreflang on a page whose canonical points elsewhere" is worth a look on your own site: the page says it is a copy of another URL and, at the same time, offers its own set of language versions. On the top 1,000 most of these are the harmless query-string cases above.
Return links
191 sites had at least one alternate I could fetch. 522 alternates in total.
| Alternate | |
|---|---|
| 200 and links back | 479 |
| 200, has hreflang, but not back to the homepage | 20 |
| 200, no hreflang at all | 8 |
| Not 200 (403, 404, 401, no answer) | 15 |
174 sites passed on every alternate I checked, 13 had at least one missing return link, and 4 couldn't be checked.
Most of the 13 have the same cause: the homepage isn't part of its own set. atlassian.com, cisco.com, ea.com and name.com list /ja/, /de-de/ or /site/us/en/index.html from the root URL but not the root URL itself, so the alternates have nothing to link back to. That's the "doesn't list itself" row above. The 8 alternates with no hreflang at all were on aliexpress.com, squarespace.com, ikea.com and blackberry.com, whose localised pages carried no hreflang in the HTML I received.
7 of the 15 non-200 alternates are on cloudflare.com: workers.dev and pages.dev land on product pages whose /es-es/, /fr-fr/ and /de-de/ versions answered 404 to me, and the /de-de/ homepage answered 403 to my user agent.
A checklist that catches all of the above
For a homepage, and any page you care about:
- One canonical, absolute, in the head. Not two, not relative.
-
Fetch the canonical URL with redirects off. It should answer 200, and its own canonical should be itself. This one check finds the trailing-slash and
www.loops, the 404 and the 403. - If you serve by location, fetch from two countries and compare the canonical.
-
hreflang: list the page itself, add
x-default, uselanguageorlanguage-REGIONwith a hyphen, and point only at canonical URLs. - Fetch every alternate and check it answers 200 and lists the others, including this page.
- Make
og:urlthe same as the canonical.
In shell, step 2 is one line per page:
curl -s -o /dev/null -w '%{http_code} %{redirect_url}\n' "$(curl -s https://example.com/ | grep -o '<link[^>]*rel="canonical"[^>]*>' | grep -o 'href="[^"]*"' | cut -d'"' -f2)"
200 and an empty redirect is what you want.
The probe, follow and analysis scripts, the raw JSON lines per domain, and the summary file that every number above comes from are kept with the date above.
If you'd like one site checked by hand, the paid GEO Mini Audit covers crawler access and machine-readability for one page you give (robots.txt per AI crawler, what your server answers each one, canonical, hreflang, meta robots and structured data) and adds paste-ready fixes (a corrected robots.txt, CDN/WAF steps for bot 403s, meta and canonical fixes, JSON-LD and llms.txt skeletons), all marked review before deploying. It queries no AI engine and measures no citations or rankings.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.