robots.txt is a plain-text file at the root of a host that tells crawlers which URL paths they may fetch. The convention dates from 1994; since September 2022 it has been an IETF standard, RFC 9309, written by Martijn Koster and three Google engineers. The RFC is short, and most robots.txt mistakes come from three rules in it that people do not expect: a crawler obeys only one group, the longest matching path decides, and a server error counts as “disallow everything”.

This guide goes through those rules with real files. Each file was parsed with two libraries on 2026-10-02: Protego 0.7.0, the parser Scrapy uses by default, and urllib.robotparser from Python 3.12, which many scripts use without knowing that it follows older rules. Where they disagree, the RFC and Google’s documentation are on Protego’s side. CPython rewrote the module for RFC 9309 in Python 3.13.14 and 3.14.5 (gh-138907), so on 2026-10-03 every file was also run with Python 3.14.7, and the text notes where the two versions differ.

Where the file lives and what it covers

A crawler requests exactly /robots.txt, in lowercase, at the top of the host (RFC 9309 §2.3). A file at https://example.com/robots.txt covers only that scheme, host and port. Google’s interpretation spells out the consequences: it does not apply to http://example.com/, https://blog.example.com/ or https://example.com:8181/, and a file at /folder/robots.txt is ignored. Every subdomain needs its own file.

The file must be UTF-8 plain text. Lines are field: value; field names are case-insensitive, while paths are case-sensitive. # starts a comment. The RFC defines only user-agent, allow and disallow; everything else, including sitemap and crawl-delay, is an extension that a crawler may or may not read. Google reads sitemap and ignores crawl-delay. Bing honors crawl-delay with values from 1 to 20 seconds, one URL per window (Bing Webmaster help).

Robots.txt is a request, not access control. RFC 9309 §3 says so directly: listing a path makes it public, and protecting it needs HTTP authentication or similar. Anyone can read your file at /robots.txt, so do not use it to hide an admin URL.

Which group a crawler obeys

A group is one or more user-agent lines followed by rules. A crawler looks for the group whose product token matches its own name (case-insensitive), merges all groups that match, and obeys only that. It falls back to the * group only when no named group matches (RFC 9309 §2.2.1). Google adds that * and a named group are never combined.

This is how a well-meant file opens the shop’s cart to Googlebot. The owner blocks /cart/ for everyone, adds a Googlebot block with no rules “to make sure Google can crawl”, and opts out of OpenAI training:

User-agent: *
Disallow: /cart/

User-agent: Googlebot
Disallow:

User-agent: GPTBot
Disallow: /

Both parsers agree on every case here. Bingbot falls into the * group and may not fetch /cart/checkout. Googlebot reads only its own group, which is empty, so /cart/checkout is allowed. Googlebot-Image also lands in that empty group: Google lists Googlebot as a second token for its image crawler (common crawlers), and both parsers matched it to the Googlebot group too. GPTBot may fetch nothing. Bing’s help page states the same rule for Bingbot: once it finds a section for itself, it ignores the generic one, so shared rules have to be repeated (Bing: How to create a robots.txt file). The fix is to copy Disallow: /cart/ into the Googlebot group, or to delete that group.

Lines other than user-agent, allow and disallow do not end a group. Google’s documentation shows a sitemap line between two user-agent lines and treats them as one group:

User-agent: a
Sitemap: https://example.com/sitemap.xml

User-agent: b
Disallow: /

Protego blocks /page for crawler a, as Google does. urllib.robotparser in Python 3.12 allows it, because it starts a new entry at the blank line; Python 3.14.7 ignores the blank line and blocks it. Blank lines carry no meaning in RFC 9309; they are only for people.

Allow, Disallow and the longest match

Within the chosen group, a crawler compares each rule’s path with the URL path from its first character. The rule with the most matching octets wins, wherever it sits in the file. When an allow and a disallow rule are equally specific, the RFC says the allow rule SHOULD win, and Google uses “the least restrictive rule” (RFC 9309 §2.2.2). If no rule matches, or the group has no rules, the URL is allowed.

Two special characters are standard: * matches any sequence of characters, and $ anchors the end of the URL. A trailing * changes nothing, because rules are prefixes already. This file was built in the ZeroTool generator with one * block, four rules and a sitemap:

User-agent: *
Disallow: /shop/
Allow: /shop/sale/
Disallow: /*?sort=
Allow: /*.css$

Sitemap: https://example.com/sitemap.xml
URL pathRule that decides (RFC 9309)Protegourllib.robotparser, Python 3.12urllib.robotparser, Python 3.14.7
/shop/sale/shoesAllow: /shop/sale/ (11 octets) beats Disallow: /shop/ (6)allowedblockedallowed
/shop/sale/shoes?sort=priceAllow: /shop/sale/ (11) beats Disallow: /*?sort= (8)allowedblockedblocked
/shop/item?sort=priceDisallow: /*?sort= (8) beats nothing longerblockedblockedblocked
/blog/?sort=newDisallow: /*?sort=blockedallowedblocked
/shop/theme.cssAllow: /*.css$ (7) beats Disallow: /shop/ (6)allowedblockedallowed
/shop/theme.css?v=2$ fails because of the query, so Disallow: /shop/blockedblockedblocked
/Shop/itemnothing matches; paths are case-sensitiveallowedallowedallowed

The second row surprises people: a sorted URL inside /shop/sale/ is crawlable, because the longer Allow beats the shorter wildcard Disallow. If you want sorted listings blocked everywhere, write a rule at least as long as the Allow, such as Disallow: /shop/sale/*?sort=.

The Python 3.12 column shows why the old urllib.robotparser is a poor test bench. It applies the first rule that matches in file order and does not understand * or $, so it treats Disallow: /*?sort= as a literal path and never matches it. Python 3.13.14, 3.14.5 and later apply the longest match and understand * and $, and Python 3.14.7 agrees with Protego on every row but the second. There it measures a wildcard rule by the part of the URL it matched: /*?sort= covers all 22 characters of /shop/sale/shoes?sort=, so it beats the 11-character Allow. Google measures “the length of the rule path” (How Google interprets the robots.txt specification), and so does Protego. If a script of yours relies on urllib.robotparser, check which Python runs it; on 3.12 or older, put each Allow above the Disallow it carves out and avoid wildcards.

Two more matching details. A rule without a trailing slash is a prefix: Disallow: /private also blocks /private-notes and /private.html. And non-ASCII paths are compared percent-encoded, so Disallow: /café and Disallow: /caf%C3%A9 are the same rule.

Missing, broken, redirected or huge files

What a crawler does when it cannot read the file matters as much as the rules. RFC 9309 §2.3.1 and Google’s documentation:

Response for /robots.txtRFC 9309Google
2xxFollow the parseable rulesSame; invalid lines are skipped. If an HTML page comes back, Google tries to extract rules from it
3xxFollow at least five redirects, then the file may be treated as unavailableFollows at least five hops, then treats it as 404
4xxUnavailable: the crawler MAY access everythingAll 4xx except 429 mean “no restrictions”
5xx, timeout, DNS errorUnreachable: assume complete disallowStops crawling for 12 hours, then uses the last good copy for up to 30 days, then behaves as if no file exists if the site is otherwise reachable
SizeParse at least 500 KiBContent after 500 KiB is ignored
CachingShould not use a copy older than 24 hours unless the file is unreachableUsually up to 24 hours; may follow Cache-Control: max-age

The 5xx row is the one that hurts in practice. A CDN or firewall rule that returns 503 for /robots.txt stops Googlebot from crawling the whole host, even though every page works in a browser. Returning 401 or 403 does the opposite: Google treats it as “no robots.txt” and crawls everything.

Blocking crawling does not block indexing

Disallowing a URL stops compliant crawlers from fetching it. It does not remove the URL from search. Google states that it may still index a disallowed URL and show it without a snippet when other pages link to it (robots.txt spec, disallow). To keep a page out of results, let it be crawled and send noindex in a meta tag or an X-Robots-Tag header. If you disallow the page as well, the crawler never sees the noindex.

The same logic applies to page resources. Google renders pages, and Apple notes that Applebot may render them too, so blocking the CSS and JavaScript a page needs can make the rendered page look broken (About Applebot).

AI crawlers and their tokens

Most AI companies now publish separate tokens for training, for their search index and for fetches a user asks for. Blocking one does not block the others:

TokenOperatorWhat disallowing it does (operator’s own documentation)
GPTBotOpenAIContent should not be used to train foundation models (OpenAI crawlers)
OAI-SearchBotOpenAISite is not shown in ChatGPT search answers, apart from navigational links
ChatGPT-UserOpenAIUser-triggered fetches; OpenAI says robots.txt rules “may not apply”
ClaudeBot, Claude-SearchBot, Claude-UserAnthropicTraining, search indexing, and user-triggered fetches; all three honor robots.txt (Anthropic help)
Google-ExtendedGoogleGemini training and grounding; no effect on Google Search inclusion or ranking. It has no user-agent string of its own (Google common crawlers)
Applebot-ExtendedAppleTraining of Apple’s foundation models; does not crawl pages and does not affect search (About Applebot)
PerplexityBot, Perplexity-UserPerplexitySearch results; user-triggered fetches “generally ignore” robots.txt (Perplexity crawlers)
CCBotCommon CrawlCommon Crawl’s open crawl, a frequent training source (CCBot)

Two details from these pages are easy to miss. Applebot follows the Googlebot group when no group names Applebot, so a Googlebot-only rule also applies to Apple. And Google-Extended and Applebot-Extended never appear in your logs, because they are control tokens read by the normal crawlers. A training opt-out file looks like this; each token needs its own block in the ZeroTool generator:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: *
Disallow:

Sitemap: https://example.com/sitemap.xml

It leaves OAI-SearchBot, Claude-SearchBot and PerplexityBot free to index the site for their search products. Robots.txt only asks; a crawler that ignores it has to be blocked at the server or CDN.

Testing a file before and after you publish

Search Console’s robots.txt report (Settings → robots.txt) lists the files Google fetched for up to 20 hosts of a property, when it fetched them, and the lines it could not parse. You can ask it to recrawl the file after a fix. It does not test URLs; for one URL, use the URL Inspection tool. Google retired the old robots.txt Tester in 2023. Bing Webmaster Tools still has a robots.txt tester that checks a URL against Bingbot or AdIdxbot rules and shows the line that blocks it.

For a check without logging in anywhere, fetch the file and run a parser that follows the RFC:

curl -sI https://example.com/robots.txt   # expect 200 and Content-Type: text/plain
pip install protego
from protego import Protego
import urllib.request

body = urllib.request.urlopen("https://example.com/robots.txt").read().decode("utf-8")
rp = Protego.parse(body)
for path in ["/shop/sale/shoes?sort=price", "/shop/item?sort=price"]:
    print(path, rp.can_fetch("https://example.com" + path, "Googlebot"))

Google also publishes the C++ parser its crawlers use, google/robotstxt, if you want the reference behavior.

Building the file with the ZeroTool generator

The robots.txt generator writes the file from blocks: pick *, Googlebot or Bingbot, or type any other token under Custom; add Allow and Disallow rows; add one Sitemap URL. All examples above came from it, and the test suite checks them against its code. Keep these behaviors in mind:

  • It opens with User-agent: * and Disallow: /, which blocks the whole site. That is right for a staging host and wrong for production, so change it first.
  • A block without rules is written as an empty Disallow:, which allows everything for that crawler, as in the Googlebot example above.
  • Each block has one user agent, and paths are copied as typed. The tool does not check that a path starts with /, warn about wildcards, or merge duplicate groups.
  • It writes one Sitemap line at the end, and no Crawl-delay.

After uploading, open https://your-host/robots.txt in a browser, check that it returns 200 as text/plain, and run the file through Protego or the Search Console report with the URLs you care about. For pages that must leave search results, the meta tag generator writes the noindex tag that robots.txt cannot replace.