Dev.to WebDev 🛠 Dev 👁 0 📖 7 min read

154 of 529 top homepages have no canonical tag. 15 point to a URL that redirects straight back.

I fetched the homepages of the Tranco top 1,000 once each and read two tags that tell crawlers which URL a page is: <link rel="canonical"> and the hreflang alternates. 529 homepages came back as real pages. 154 of the 52

I fetched the homepages of the Tranco top 1,000 once each and read two tags that tell crawlers which URL a page is: <link rel="canonical"> and the hreflang alternates. 529 homepages came back as real pages. 154 of the 529 have no canonical tag at all. Of the 375 that do, 317 point at the URL they were served from and 58 point somewhere else. I then fetched those 58 targets: 20 redirect, 15 of them straight back to the page that named them, one is a 404 and one is a 403.

The hreflang side is in better shape than its reputation. 197 homepages declare language versions. I fetched up to three alternates for each and checked whether they link back. 479 of the 522 alternates I fetched do; 15 of the rest didn't answer 200.

Below: the limits first, the numbers, the mistakes that repeat, and a short checklist.

Limits first (measured 2026-10-06/07)

  • One fetch per domain, from Japan. Homepages on 2026-10-06 between 21:30 and 21:54 UTC, then the canonical targets and alternates between 21:55 and 22:03 UTC, one request at a time, 10 second timeout, redirects followed, no JavaScript. Several sites serve a Japanese page to a Japanese IP, and their canonical changes with it. From another country you'd see other URLs for some of these.
  • A bot user agent. I identified myself (glovrex-measure/1.0). 96 domains answered 403 to that, which is fewer pages than my earlier runs with a browser user agent. Those sites are unknown here, not "no canonical".
  • List and set: Tranco list 56WKN, top 1,000. Many entries are CDN, API or tracking domains with no homepage: 314 gave no HTTP answer (DNS, TLS, timeout), 148 answered with something other than 200, 9 answered 200 with no HTML head or an "Access Denied" page. n = 529 below.
  • Homepage only, <head> only. I read canonical and hreflang from the HTML head and from the HTTP Link header. hreflang in an XML sitemap isn't covered, so a site can be fine there and still show up as "no hreflang" here.
  • Return links are sampled. At most three alternates per site, skipping x-default and the page itself.
  • I don't know how AI crawlers use these tags. Google and Bing document canonical and hreflang. OpenAI, Anthropic and Perplexity don't say whether their crawlers read them. Answer engines that build on a search index inherit whatever the index decided.

The canonical numbers

n = 529 homepages.

Homepages
No canonical (head or Link header) 154
Canonical = the URL served 317
Canonical points elsewhere 58
More than one canonical tag 4
Two canonicals that disagree 1
Relative canonical (/, /us/) 2
og:url present 291
og:url disagrees with canonical 19

The 58 that point elsewhere, by what differs:

Difference Homepages
Only the query string 18
A different host 15
A different path 14
Only a trailing slash 8
Only www. 3

The query-string group is mostly fine. bing.com, ya.ru and youtu.be redirect with a tracking or redirect parameter and the canonical drops it, which is what a canonical is for.

What the 58 targets answered

Canonical target
200, and declares itself canonical 32
Redirects to another URL 20
No answer 2
404 1
403 1
200, no canonical of its own 1
200, canonical points on again 1

The 20 redirects are the interesting group. 15 of them redirect to exactly the URL that declared them, 3 more come back to it with a fresh query parameter (ya.ru twice, salesforce.com), and 2 go somewhere new (amp.dev, hihonorcloud.com). The loops look like this:

  • Trailing slash, both ways. huawei.com serves /en/ and names /en as canonical, and /en redirects to /en/. agora.io does the same. ubisoft.com does it in the other direction.
  • www. both ways. go.com is served at www.go.com with canonical go.com, which redirects to www.go.com. bugsnag.com does the same.
  • Mobile and edition hosts. vk.com sends my user agent to m.vk.com, whose canonical is vk.com, which redirects back to m.vk.com. cnn.com lands on edition.cnn.com with canonical www.cnn.com.
  • A path the server won't serve. epa.gov names /home, which redirects to /.

A canonical that redirects isn't fatal. Google calls rel=canonical and redirects both signals and picks one. But here the two signals point in opposite directions, so the site has handed the choice to the crawler.

The two hard failures are a stanford.edu homepage whose canonical is https://www.stanford.edu/home, which answers 404 (rechecked with curl, same result), and synology.com, whose canonical is https://www.synology.com// with a double slash, which answers 403.

The canonical depends on where you are

Six of the "different host" and "different path" cases, from four companies, point at a regional page when seen from Japan:

  • facebook.com (and fb.com, fbcdn.net): canonical https://ja-jp.facebook.com/
  • android.com: canonical https://www.android.com/intl/ja_jp/
  • forbes.com: canonical https://www.forbes.com/home_asia/
  • intel.com: served from www.intel.co.jp, canonical on www.intel.com

This can be intentional. It also means a crawler in the US and a crawler in Japan are told different things about the same URL, and a homepage can quietly be declared a copy of a regional page. If you serve by location, check the canonical from more than one country.

og:url versus canonical

19 of 291 disagree. Most are small (instagram.com has canonical www.instagram.com and og:url instagram.com), but two are real: roblox.com's homepage says og:url is /CreateAccount, and gandi.net's canonical is /en-US while og:url is /en. Link previews and some crawlers read og:url, so it's worth keeping it in sync.

The hreflang numbers

n = 197 homepages with hreflang (37% of 529). Median 13 alternates, maximum 270.

Homepages
Includes x-default 137
Lists itself among its alternates 183
Doesn't list itself 14
hreflang on a page whose canonical points elsewhere 19
A code that isn't a valid language or region 7
A 3-letter language code (fil, ceb, skr) 5
A withdrawn code (iw, in) 4
Same code, two different URLs 4
Relative or non-HTTP URL 2

The 7 invalid ones: weebly.com and amp.dev use underscores (en_GB, pt_BR); viber.com uses ua and gr, which are country codes, not languages; aliyun.com uses tc; hp.com writes country first (py-es, us-en) plus internal names like lamerica_nsc_carib-en; ezviz uses region eu; shein.com uses hr-eur and zh-tw-sg. Google's localized versions guide asks for an ISO 639-1 language with an optional ISO 3166-1 region. I count fil or iw separately because they are valid language tags, but fil has no ISO 639-1 code and iw was replaced by he in 1989.

"hreflang on a page whose canonical points elsewhere" is worth a look on your own site: the page says it is a copy of another URL and, at the same time, offers its own set of language versions. On the top 1,000 most of these are the harmless query-string cases above.

Return links

191 sites had at least one alternate I could fetch. 522 alternates in total.

Alternate
200 and links back 479
200, has hreflang, but not back to the homepage 20
200, no hreflang at all 8
Not 200 (403, 404, 401, no answer) 15

174 sites passed on every alternate I checked, 13 had at least one missing return link, and 4 couldn't be checked.

Most of the 13 have the same cause: the homepage isn't part of its own set. atlassian.com, cisco.com, ea.com and name.com list /ja/, /de-de/ or /site/us/en/index.html from the root URL but not the root URL itself, so the alternates have nothing to link back to. That's the "doesn't list itself" row above. The 8 alternates with no hreflang at all were on aliexpress.com, squarespace.com, ikea.com and blackberry.com, whose localised pages carried no hreflang in the HTML I received.

7 of the 15 non-200 alternates are on cloudflare.com: workers.dev and pages.dev land on product pages whose /es-es/, /fr-fr/ and /de-de/ versions answered 404 to me, and the /de-de/ homepage answered 403 to my user agent.

A checklist that catches all of the above

For a homepage, and any page you care about:

  1. One canonical, absolute, in the head. Not two, not relative.
  2. Fetch the canonical URL with redirects off. It should answer 200, and its own canonical should be itself. This one check finds the trailing-slash and www. loops, the 404 and the 403.
  3. If you serve by location, fetch from two countries and compare the canonical.
  4. hreflang: list the page itself, add x-default, use language or language-REGION with a hyphen, and point only at canonical URLs.
  5. Fetch every alternate and check it answers 200 and lists the others, including this page.
  6. Make og:url the same as the canonical.

In shell, step 2 is one line per page:

curl -s -o /dev/null -w '%{http_code} %{redirect_url}\n' "$(curl -s https://example.com/ | grep -o '<link[^>]*rel="canonical"[^>]*>' | grep -o 'href="[^"]*"' | cut -d'"' -f2)"

200 and an empty redirect is what you want.

The probe, follow and analysis scripts, the raw JSON lines per domain, and the summary file that every number above comes from are kept with the date above.

If you'd like one site checked by hand, the paid GEO Mini Audit covers crawler access and machine-readability for one page you give (robots.txt per AI crawler, what your server answers each one, canonical, hreflang, meta robots and structured data) and adds paste-ready fixes (a corrected robots.txt, CDN/WAF steps for bot 403s, meta and canonical fixes, JSON-LD and llms.txt skeletons), all marked review before deploying. It queries no AI engine and measures no citations or rankings.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.