Dev.to WebDev 🛠 Dev 👁 0 📖 4 min read

Google Search Console for Developers: Why Your Fast, Clean Site Is Only Partly Indexed

Shipping a page and getting it into Google are two separate deployments, and only one of them has a CI pipeline. Picture a small team that launches a documentation site. It has 4,000 URLs, a tidy sitemap, and a Lighthou

Google Search Console for Developers: Why Your Fast, Clean Site Is Only Partly Indexed

Shipping a page and getting it into Google are two separate deployments, and only one of them has a CI pipeline.

Picture a small team that launches a documentation site. It has 4,000 URLs, a tidy sitemap, and a Lighthouse score everyone screenshots. Three weeks later, search traffic is flat.

This is a made-up team, but you may know the feeling. The build is green. The index is not.

Start with the Page indexing report, not the Performance report

Performance tells you how visible you are. The Page indexing report tells you whether you are eligible to be visible at all. Google's documentation states the goal plainly: get the canonical version of every important page indexed, while duplicates and alternates stay out.

In our example, 1,300 of 4,000 known URLs are indexed. The rest sit in reasons you can read as a to-do list.

Illustrative page indexing report: 4,000 known URLs grouped by reason. The numbers are made up.

Two reasons look alike and are not:

  • Crawled, currently not indexed. Google fetched the page and chose not to index it. Google says you do not need to resubmit it. Treat it as a content or quality question, not a crawling one.
  • Discovered, currently not indexed. Google knows the URL but has not crawled it. Google's explanation is that the crawl was expected to overload the site, so it was rescheduled. That points at your server and your URL count.

Different reasons, different fixes. Do not paste the same "request indexing" click on both.

Use URL Inspection as a debugger

Paste a URL into the inspection bar and you get Google's indexed version of the page. That is not a live test. For the current page, run the Live Test, and open View tested page to see the rendered result, loaded resources and JavaScript output.

Two details save hours. First, the Google-selected canonical field shows what Google chose, and per the docs you can read it only from indexed data, because the live test cannot predict it. Second, a passing live test does not guarantee a place in search results. It does not check quality guidelines, manual actions or removals.

If your page says canonical A and Google picked canonical B, you have a duplicate-content problem or a redirect problem. Look at "Duplicate without user-selected canonical": Google picked a canonical for you because you did not declare one. Declare it.

Read Crawl Stats like server logs you did not have to parse

The Crawl Stats report covers total requests, bytes downloaded, average response time, host status, crawl responses and crawl purpose, according to the Crawl Stats help page.

The detail developers miss: every hop of a redirect chain counts as a separate request. A chain of three redirects spends three crawl requests to reach one page.

Illustrative Crawl Stats: after a deploy, average response time climbs and daily crawl requests fall. Data is made up.

Watch the host status too. Google throttles when robots.txt fails to return a valid file or a 404, and when the server stops answering. A slow deploy that doubles response time can quietly shrink how much of your site gets fetched.

Fix what you can measure: Core Web Vitals and rich results

Google's targets are LCP within 2.5 seconds, INP under 200 milliseconds and CLS under 0.1, as listed on its Core Web Vitals page. Search Console groups URLs by status, so you fix a template once and clear a whole group.

Validation is a loop, not a button. When you start validation, Google checks instances and reports states like Looking good or Passing as it confirms the fix.

Illustrative field data before and after a template fix, against Google's

Structured data errors work the same way. Fix the template, inspect a sample URL, then validate the group.

Stop clicking, start querying

The Search Console UI caps what you can browse. The Search Analytics API lets you page through up to 25,000 rows per request with rowLimit and startRow. Set dataState to all for fresh data, and expect the latest days to change. The response metadata flags the first incomplete date.

A small script that pulls page and query pairs weekly gives you a trend line the dashboard never will. Put it in a repo. Review it like any other monitoring job.

A weekly routine that fits in a sprint

  1. Open Page indexing and read the top three non-indexed reasons.
  2. Inspect one URL from each and decide: content, canonical or crawl budget.
  3. Check Crawl Stats for response time and host status after each deploy.
  4. Validate one fix at a time so you can tell what worked.
  5. Pull the API data and compare equal periods.

Our imaginary team does not need a rewrite. It needs a feedback loop between what it ships and what Google accepts.

If you want more posts on building sites that bring in clients organically, follow Orin Green on DEV. The next one starts with the page everyone built and nobody found.

Sources: Google Search Console Help (Page indexing report, URL Inspection tool, Crawl Stats report), Google Search Central (Core Web Vitals), Google Search Console API reference (searchanalytics.query). Charts use made-up illustrative data.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.