Robots.txt Best Practices for SEO in 2026
Robots.txt tells search engine crawlers which parts of your site to visit. Used correctly, it protects crawl budget. Used incorrectly, it can hide your content from Google. How Robots.txt Works Located at: you
Robots.txt tells search engine crawlers which parts of your site to visit.
Used correctly, it protects crawl budget. Used incorrectly, it can hide your content from Google.
How Robots.txt Works
Located at: yourdomain.com/robots.txt
Syntax: User-agent + Disallow/Allow rules.
User-agent: *
Disallow: /admin/
Disallow: /staging/
Allow: /
- User-agent: * = applies to all crawlers
- Disallow: /admin/ = blocks crawlers from that directory
- Allow: / = permits crawling everything not explicitly blocked
What to Block with Robots.txt
Safe to block:
- /admin/ (backend login pages)
- /staging/ (test versions of pages)
- /search/ (internal search result pages create URL variations)
- /cart/ and /checkout/ (no indexing value)
- URL parameters that create duplicate content
Never block with robots.txt:
- Your main content pages
- CSS and JS files Google needs to render your pages
- Anything linked from your sitemap
Common Mistake: Blocking CSS/JS
Google needs CSS and JS to render your pages as users see them.
Blocking these files: Google sees a broken page. Rankings suffer.
Test: GSC URL Inspection → View crawled page → Does it look correct?
Robots.txt vs. Noindex
Robots.txt blocks crawling. Noindex blocks indexing.
If you disallow a page in robots.txt: Google cannot crawl it to see the noindex tag.
Use noindex for pages you want Google to visit but not index.
Use robots.txt for pages you want Google to not visit at all.
Robots.txt and technical SEO optimization: yositeup.com
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.