Dev.to WebDev 🛠 Dev 👁 0 📖 5 min read

Robots.txt Deep Dive: What It Can and Cannot Do

Robots.txt is one of the smallest files on your site and one of the most misunderstood. It's a plain text file, and a single wrong line can hide your whole site from search engines. Even more often, people expect it to d

Robots.txt is one of the smallest files on your site and one of the most misunderstood. It's a plain text file, and a single wrong line can hide your whole site from search engines. Even more often, people expect it to do something it was never built for.

Let's sort out what it actually does.

What robots.txt controls

Robots.txt controls crawling. That's all. It tells compliant bots which URLs they may or may not request.

  • It lives at the root of a host: https://example.com/robots.txt
  • Each host, protocol and port needs its own file (blog.example.com doesn't inherit from example.com)
  • It's a polite request, not a lock. Reputable crawlers like Googlebot and Bingbot follow it. Malicious bots ignore it
  • It's public. Anyone can read it, so never use it to hide private pages

It does not control indexing, and it does not provide security.

The basic syntax

A robots.txt file is made of groups. Each group starts with a user-agent and is followed by rules.

User-agent: *
Disallow: /admin/
Allow: /admin/help/

Sitemap: https://example.com/sitemap.xml

User-agent

This names the bot a group applies to.

  • * means all bots
  • Googlebot, Bingbot and other names target specific crawlers
  • A bot follows only the most specific group that matches it. It does not combine that group with the * group

That last point trips people up. If you write a Googlebot group, Googlebot ignores the * group completely, so repeat any shared rules inside it.

Disallow

Blocks a path from being crawled.

  • Disallow: /private/ blocks everything under that folder
  • Disallow: / blocks the entire site
  • Disallow: (empty) blocks nothing

Allow

Makes an exception inside a blocked area.

  • Useful for opening one subfolder inside a blocked directory
  • When rules conflict, Google follows the most specific (longest) match. If it's a tie, the less restrictive rule wins

Sitemap

Points crawlers to your XML sitemap.

  • Use a full absolute URL
  • It isn't tied to any user-agent group, so it can sit anywhere in the file
  • You can list more than one

Wildcards

Google and Bing support two special characters:

  • * matches any sequence of characters
  • $ marks the end of a URL
User-agent: *
Disallow: /*.pdf$
Disallow: /search*

The first rule blocks URLs ending in .pdf. Without the $, it would also block something like /file.pdf?download=1.

Directory blocking

Be careful with trailing slashes. They change the meaning.

  • Disallow: /blog/ blocks the folder and everything inside it
  • Disallow: /blog blocks anything that starts with /blog, including /blog-news and /blogger-tools

Paths are also case-sensitive. /Admin/ and /admin/ are different things.

Query parameters

Parameters can create thousands of near-identical URLs: sorting, filters, session IDs, tracking tags. Robots.txt is a common way to stop crawlers from wasting time on them.

User-agent: *
Disallow: /*?sort=
Disallow: /*?*sessionid=

A few cautions:

  • Block only parameters that truly produce duplicate or useless pages
  • If a parameter changes the actual content (like ?page=2 on a real listing), blocking it can hide products or articles from crawlers
  • For duplicates you want consolidated rather than ignored, a canonical tag is often the better tool

Bot-specific rules

You can give different bots different instructions.

User-agent: Googlebot
Disallow: /internal-search/

User-agent: GPTBot
Disallow: /

User-agent: *
Disallow: /admin/
  • This is how people opt out of certain AI crawlers while staying open to search engines
  • Remember that each bot reads only its own group
  • Crawl-delay is ignored by Google. Bing supports it. Don't rely on it for Googlebot

Why robots.txt is not an indexing directive

This is the big misunderstanding.

  • Disallow stops Google from fetching a page
  • It doesn't stop Google from knowing the URL exists
  • If other pages link to a blocked URL, Google can still index the bare address and show it in results with no description

So a blocked page can still appear in search. It just appears without content, because Google was never allowed to read it.

Google also no longer supports a noindex line inside robots.txt. Any guide telling you to write Noindex: /page/ in the file is out of date.

Blocking crawling vs noindex

These two tools solve different problems.

  • Disallow in robots.txt: "Don't fetch this URL"
  • noindex (meta tag or X-Robots-Tag header): "You may fetch this, but don't keep it in the index"
<meta name="robots" content="noindex">

The trap is combining them. If you block a page in robots.txt and add a noindex tag, Google never fetches the page, so it never sees the noindex. The URL can stay indexed for a long time.

To remove a page from search properly:

  1. Let Google crawl it (no Disallow)
  2. Serve a noindex tag or header
  3. Wait for Google to recrawl and drop it
  4. Only then consider blocking it, if you still want to

For pages that are truly private, use authentication. Neither robots.txt nor noindex is real protection.

Common syntax mistakes

  • Leaving Disallow: / live after launch. The classic staging-site disaster. Check it after every migration
  • Blocking CSS and JavaScript. Google needs these to render your pages properly
  • Wrong case or missing slash. /Blog/ won't match /blog/
  • Putting rules before any User-agent line. Rules must belong to a group
  • Relative sitemap URLs. Use the full https:// address
  • Using robots.txt to hide sensitive content. The file is public, so it can actually point people to what you want hidden
  • Assuming * rules apply to every group. They don't, as covered above
  • Wrong location or format. It must sit at the root, be plain text, and return a 200 status. If it returns a server error for a long time, Google may treat the site cautiously

A sensible starter file

User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /*?sessionid=

Sitemap: https://example.com/sitemap.xml

Short, clear and easy to audit. Add rules only when you have a reason, and write a comment (# reason) so future you remembers why.

How to test it

  • Use the robots.txt report in Google Search Console to see how Google reads your file
  • Run the URL Inspection tool on important pages and confirm they aren't blocked
  • Open /robots.txt in your browser after every deployment

Final thoughts

Think of robots.txt as a traffic sign for crawlers, not a vault door and not a deletion button. It manages where bots spend their time. If you want something kept out of search results, use noindex. If you want something truly private, use a login.

Have you ever found a rogue Disallow: / in production? Tell me about it in the comments.

Also , if you're interested, do visit my website - anoshbb.com

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.