Robots.txt Generator
Generate a valid robots.txt file instantly. Configure User-agent rules, Allow/Disallow paths, and Sitemap URL — no sign-up required.
- Runs in your browser
- Your data never leaves your browser
- Free · No Sign-Up
Scan with WeChat to share this tool
Examples, details and FAQ Worked examples, how it compares with other tools, and answers to common questions.
Example: a Typical Site File
Keep the * block, change its rule to Disallow /admin/ and add Allow /admin/public/; add a second block with the custom name GPTBot and Disallow /; enter the sitemap URL. The output is:
User-agent: *
Disallow: /admin/
Allow: /admin/public/
User-agent: GPTBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
Under RFC 9309 the longest matching path wins, whatever the line order, and when an Allow and a Disallow match with the same length the Allow should win. So /admin/public/help may be crawled, /admin/users may not, and GPTBot may fetch nothing. Scrapy’s Protego parser returned exactly that. Python 3.12’s urllib.robotparser did not: it applies the first matching line in file order and reported /admin/public/help as blocked. CPython rewrote the module for RFC 9309 in Python 3.13.14 and 3.14.5 (gh-138907), and Python 3.14.7 allows that URL. If one of your own scripts uses it on an older Python, put the Allow line above the Disallow line; RFC-compliant crawlers read both orders the same way.
Example: Allow One Crawler Only
User-agent: *
Disallow: /
User-agent: Googlebot
Disallow:
Build it with two blocks: the default * block with Disallow /, and a Googlebot block with no rules. A block without rules gets an empty Disallow: line, which allows everything. Googlebot reads only its own group, so the * rule does not apply to it.
Limits and Things robots.txt Does Not Do
- No validation. Paths are trimmed and otherwise copied as typed. RFC 9309 paths start with
/; a value such asadmin/is not a valid rule. Empty paths are written as empty rules. - One Sitemap line. For several sitemaps, copy the line and add the others by hand.
Crawl-delayand other non-standard lines are not generated; Google ignoresCrawl-delay. - Wildcards
*and$are part of RFC 9309 and can be typed into a path, for exampleDisallow: /*?sessionid=; the tool does not check them. - Blocking is not removal. A disallowed URL can still appear in search results without a description if other pages link to it. To keep a page out of results, let it be crawled and send
noindex, which the meta tag generator can write. - Size and errors. RFC 9309 asks crawlers to parse at least the first 500 KiB. If the file returns a 4xx status, crawlers may crawl everything; if it returns 5xx or cannot be reached, they should assume everything is disallowed.
Common robots.txt Patterns
- Block all bots:
User-agent: *+Disallow: / - Allow all bots:
User-agent: *+Disallow:(empty value) - Block a specific directory:
Disallow: /admin/ - Allow Googlebot only: Block
*, then add a separate Googlebot block withDisallow:
After generating your file, upload it to your web server root so it is accessible at
https://yourdomain.com/robots.txt. Each host and protocol needs its own file: rules at example.com do not cover blog.example.com.
FAQ
What is a robots.txt file?
robots.txt is a plain-text file placed at the root of a website (e.g. https://example.com/robots.txt) that tells web crawlers which pages or directories they are allowed or not allowed to access. It follows the Robots Exclusion Protocol, standardized as RFC 9309.
Does robots.txt block all bots?
No. robots.txt is a convention, not a security measure. Well-behaved crawlers like Googlebot and Bingbot respect it, but malicious bots may ignore it entirely. Use server-level access controls to truly restrict content.
What does 'Disallow: /' mean?
It tells the specified User-agent to not crawl any page on the site. Using 'User-agent: *' combined with 'Disallow: /' blocks all compliant bots from the entire site. This is also the tool's starting state, so change it before you publish the file.
Can I have multiple User-agent blocks?
Yes. Click '+ Add User-agent Block' to add more. A crawler follows only the most specific group that names it, so a Googlebot block replaces the * block for Googlebot instead of adding to it; repeat any shared rules in both blocks.
Where should I put the Sitemap directive?
Anywhere in the file. Sitemap lines are not part of any User-agent group, so their position does not matter; the tool writes the line at the end. It tells crawlers where your XML sitemap is, which helps them discover the pages you want crawled.
Are the paths and URLs I type sent anywhere or saved?
No. The file is built in your browser tab. The tool does not fetch the Sitemap URL, and it does not send your paths or bot names to a server or write them to browser storage, so reloading the page brings back the starting block. When you commit a change (change a field and leave it, pick a bot, add or remove a rule or block), the site's analytics records a usage event with the tool name and the action, not what you typed.