Robots.txt Deep Dive: What It Can and Cannot Do
Robots.txt is one of the smallest files on your site and one of the most misunderstood. It's a plain text file, and a single wrong line can hide your whole site from search engines. Even more often, people expect it to d
Robots.txt is one of the smallest files on your site and one of the most misunderstood. It's a plain text file, and a single wrong line can hide your whole site from search engines. Even more often, people expect it to do something it was never built for.
Let's sort out what it actually does.
What robots.txt controls
Robots.txt controls crawling. That's all. It tells compliant bots which URLs they may or may not request.
- It lives at the root of a host:
https://example.com/robots.txt - Each host, protocol and port needs its own file (
blog.example.comdoesn't inherit fromexample.com) - It's a polite request, not a lock. Reputable crawlers like Googlebot and Bingbot follow it. Malicious bots ignore it
- It's public. Anyone can read it, so never use it to hide private pages
It does not control indexing, and it does not provide security.
The basic syntax
A robots.txt file is made of groups. Each group starts with a user-agent and is followed by rules.
User-agent: *
Disallow: /admin/
Allow: /admin/help/
Sitemap: https://example.com/sitemap.xml
User-agent
This names the bot a group applies to.
-
*means all bots -
Googlebot,Bingbotand other names target specific crawlers - A bot follows only the most specific group that matches it. It does not combine that group with the
*group
That last point trips people up. If you write a Googlebot group, Googlebot ignores the * group completely, so repeat any shared rules inside it.
Disallow
Blocks a path from being crawled.
-
Disallow: /private/blocks everything under that folder -
Disallow: /blocks the entire site -
Disallow:(empty) blocks nothing
Allow
Makes an exception inside a blocked area.
- Useful for opening one subfolder inside a blocked directory
- When rules conflict, Google follows the most specific (longest) match. If it's a tie, the less restrictive rule wins
Sitemap
Points crawlers to your XML sitemap.
- Use a full absolute URL
- It isn't tied to any user-agent group, so it can sit anywhere in the file
- You can list more than one
Wildcards
Google and Bing support two special characters:
-
*matches any sequence of characters -
$marks the end of a URL
User-agent: *
Disallow: /*.pdf$
Disallow: /search*
The first rule blocks URLs ending in .pdf. Without the $, it would also block something like /file.pdf?download=1.
Directory blocking
Be careful with trailing slashes. They change the meaning.
-
Disallow: /blog/blocks the folder and everything inside it -
Disallow: /blogblocks anything that starts with/blog, including/blog-newsand/blogger-tools
Paths are also case-sensitive. /Admin/ and /admin/ are different things.
Query parameters
Parameters can create thousands of near-identical URLs: sorting, filters, session IDs, tracking tags. Robots.txt is a common way to stop crawlers from wasting time on them.
User-agent: *
Disallow: /*?sort=
Disallow: /*?*sessionid=
A few cautions:
- Block only parameters that truly produce duplicate or useless pages
- If a parameter changes the actual content (like
?page=2on a real listing), blocking it can hide products or articles from crawlers - For duplicates you want consolidated rather than ignored, a canonical tag is often the better tool
Bot-specific rules
You can give different bots different instructions.
User-agent: Googlebot
Disallow: /internal-search/
User-agent: GPTBot
Disallow: /
User-agent: *
Disallow: /admin/
- This is how people opt out of certain AI crawlers while staying open to search engines
- Remember that each bot reads only its own group
-
Crawl-delayis ignored by Google. Bing supports it. Don't rely on it for Googlebot
Why robots.txt is not an indexing directive
This is the big misunderstanding.
-
Disallowstops Google from fetching a page - It doesn't stop Google from knowing the URL exists
- If other pages link to a blocked URL, Google can still index the bare address and show it in results with no description
So a blocked page can still appear in search. It just appears without content, because Google was never allowed to read it.
Google also no longer supports a noindex line inside robots.txt. Any guide telling you to write Noindex: /page/ in the file is out of date.
Blocking crawling vs noindex
These two tools solve different problems.
- Disallow in robots.txt: "Don't fetch this URL"
-
noindex (meta tag or
X-Robots-Tagheader): "You may fetch this, but don't keep it in the index"
<meta name="robots" content="noindex">
The trap is combining them. If you block a page in robots.txt and add a noindex tag, Google never fetches the page, so it never sees the noindex. The URL can stay indexed for a long time.
To remove a page from search properly:
- Let Google crawl it (no
Disallow) - Serve a
noindextag or header - Wait for Google to recrawl and drop it
- Only then consider blocking it, if you still want to
For pages that are truly private, use authentication. Neither robots.txt nor noindex is real protection.
Common syntax mistakes
-
Leaving
Disallow: /live after launch. The classic staging-site disaster. Check it after every migration - Blocking CSS and JavaScript. Google needs these to render your pages properly
-
Wrong case or missing slash.
/Blog/won't match/blog/ -
Putting rules before any
User-agentline. Rules must belong to a group -
Relative sitemap URLs. Use the full
https://address - Using robots.txt to hide sensitive content. The file is public, so it can actually point people to what you want hidden
-
Assuming
*rules apply to every group. They don't, as covered above - Wrong location or format. It must sit at the root, be plain text, and return a 200 status. If it returns a server error for a long time, Google may treat the site cautiously
A sensible starter file
User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /*?sessionid=
Sitemap: https://example.com/sitemap.xml
Short, clear and easy to audit. Add rules only when you have a reason, and write a comment (# reason) so future you remembers why.
How to test it
- Use the robots.txt report in Google Search Console to see how Google reads your file
- Run the URL Inspection tool on important pages and confirm they aren't blocked
- Open
/robots.txtin your browser after every deployment
Final thoughts
Think of robots.txt as a traffic sign for crawlers, not a vault door and not a deletion button. It manages where bots spend their time. If you want something kept out of search results, use noindex. If you want something truly private, use a login.
Have you ever found a rogue Disallow: / in production? Tell me about it in the comments.
Also , if you're interested, do visit my website - anoshbb.com
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.