Dev.to WebDev 🛠 Dev 👁 0 📖 4 min read

Is my site down for everyone, or just me? Here is how to actually find out

You cannot answer this question from your own browser, and that is not a figure of speech — it is a structural property of your session. You are logged in. You have cookies, a role, permissions, drafts nobody else can s

You cannot answer this question from your own browser, and that is not a figure of speech — it is a structural property of your session.

You are logged in. You have cookies, a role, permissions, drafts nobody else can see, a service worker holding a cached copy, and a CDN edge that already warmed up for your IP. When you load your own page, the system shows you the version it shows you. That version is almost always the flattering one.

This week the API told me a repository was "visibility": "public" at the exact moment every logged-out visitor on earth was getting a 404. The dashboard was not lying. It was answering a different question from the one I thought I was asking.

Here is how to ask the right one.

The one-line version

curl -sI -o /dev/null -w '%{http_code}\n' https://yoursite.example.com

No cookies, no session, no extensions. That is closer to the truth than anything your browser will tell you — but it is still not enough, for three reasons below.

Why the quick check is not enough

1. A 200 does not mean the page exists. Plenty of sites serve their error template with a 200 status. A shop returns the storefront instead of the product you deleted. A framework's router falls back to the index page. Uptime monitors call all of that healthy, and so does curl -I.

This is the failure mode that catches people, because everything downstream — your monitor, your status page, your CI check — agrees the page is fine.

2. HEAD and GET do not always agree. Some servers, CDNs and WAFs handle them differently. Check with the method a reader actually uses.

3. Your own links are not checked at all. The dangerous ones are the links inside pages you already published: in your docs, in your README, in a PDF somebody downloaded last year. Those keep pointing at things you renamed, and nothing you own will ever tell you.

What a real check looks like

Fetch it the way a stranger does, then judge the response, not the status code:

const res = await fetch(url, {
  redirect: 'follow',
  credentials: 'omit',        // no cookies, ever
  cache: 'no-store',
  headers: { 'User-Agent': UA, Accept: 'text/html,*/*' },
});

Then ask seven questions of what comes back:

Question Why it matters
Is it a real 404 / 401 / 403? Your session sees the page. Nobody else does.
Is it a soft 404 — 200 with an error page? Monitors call this healthy. Deleted pages do it constantly.
Did the redirect change host or path? The link you printed is not the page they land on.
Is there a noindex you did not intend? Perfectly live, permanently invisible.
Is the body empty without JavaScript? That is what Google, Slack and every link preview see.
Do the links on the page still resolve? The ones that travel inside downloaded files are the worst kind.
Does it differ from what you see logged in? If yes, that difference is your bug.

Detecting the soft 404 is the fiddly part, and the trap is the obvious implementation. My first attempt searched the whole document for phrases like page not found and promptly flagged my own profile — it found the words inside an article that was about 404 errors.

Key off the <title> first, and only fall back to body text on genuinely short pages:

if (SOFT_404.some((re) => re.test(title))) {
  return { level: 'FAIL', note: `soft 404: replies 200 but the title says "${title}"` };
}
if (text.length < 400 && SOFT_404.some((re) => re.test(text.slice(0, 300)))) {
  return { level: 'FAIL', note: 'soft 404: replies 200 with an error page' };
}

And measure visible text from <body> only. The <head> of a modern site is tens of kilobytes of inlined CSS and preloads; twenty kilobytes in, a real page has not started yet. Any "is this empty" heuristic that includes the head will be wrong about every site built after about 2015.

A checker that cries wolf stops being read, and then it stops working. That is the whole design constraint.

The script

I packaged the above. One file, no dependencies, Node 18+, MIT.

curl -s https://files.catbox.moe/t97937.js -o outsidein.js

node outsidein.js https://yoursite.example.com
node outsidein.js --links https://yoursite.example.com   # also every link on the page
node outsidein.js urls.txt --json                        # for CI
  OK   200  https://example.com/product
             Your product page title

 FAIL  404  https://example.com/old-bundle
             -> not found for the public
             linked from https://example.com/

 WARN  200  https://example.com/app
             -> empty without JavaScript - a crawler sees nothing here

7 checked, 1 broken, 2 worth a look, all of it without a session.

It exits non-zero, so it goes straight into a release script:

node outsidein.js urls.txt || exit 1

Keep a urls.txt of everything you have ever published — product pages, docs, articles with links inside them, the URL you printed on something physical. Run it after every deploy.

The habit, which matters more than the script

If the check runs as the actor, it is not a check.

That generalises well past uptime. Publishing flows, permission changes, share links, paywalls, feature flags, "did that email actually send" — all of them look correct from inside the session that configured them, because that session is the one privileged view where everything is already true.

Log out. Use a different client. Ask a stranger. It takes thirty seconds and it is the only version of the answer that is worth anything.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.