Dev.to WebDev 🛠 Dev 👁 0 📖 1 min read

Robots.txt Best Practices for SEO in 2026

Robots.txt tells search engine crawlers which parts of your site to visit. Used correctly, it protects crawl budget. Used incorrectly, it can hide your content from Google. How Robots.txt Works Located at: you

Robots.txt tells search engine crawlers which parts of your site to visit.
Used correctly, it protects crawl budget. Used incorrectly, it can hide your content from Google.

How Robots.txt Works

Located at: yourdomain.com/robots.txt
Syntax: User-agent + Disallow/Allow rules.

User-agent: *
Disallow: /admin/
Disallow: /staging/
Allow: /
  • User-agent: * = applies to all crawlers
  • Disallow: /admin/ = blocks crawlers from that directory
  • Allow: / = permits crawling everything not explicitly blocked

What to Block with Robots.txt

Safe to block:

  • /admin/ (backend login pages)
  • /staging/ (test versions of pages)
  • /search/ (internal search result pages create URL variations)
  • /cart/ and /checkout/ (no indexing value)
  • URL parameters that create duplicate content

Never block with robots.txt:

  • Your main content pages
  • CSS and JS files Google needs to render your pages
  • Anything linked from your sitemap

Common Mistake: Blocking CSS/JS

Google needs CSS and JS to render your pages as users see them.
Blocking these files: Google sees a broken page. Rankings suffer.
Test: GSC URL Inspection → View crawled page → Does it look correct?

Robots.txt vs. Noindex

Robots.txt blocks crawling. Noindex blocks indexing.
If you disallow a page in robots.txt: Google cannot crawl it to see the noindex tag.
Use noindex for pages you want Google to visit but not index.
Use robots.txt for pages you want Google to not visit at all.

Robots.txt and technical SEO optimization: yositeup.com

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.