Robots.txt Generator

Generate a valid robots.txt file instantly. Configure User-agent rules, Allow/Disallow paths, and Sitemap URL — no sign-up required.

  • Runs in your browser
  • Your data never leaves your browser
  • Free · No Sign-Up
Add a separate User-agent block. Each block has its own bot selection and Allow or Disallow rules; deleting a block removes it from the output.
Copy the complete generated file. Input changes update it immediately. Ctrl/⌘+L clears text, the Sitemap URL and the result while keeping blocks, rule types and User-agent selections. Retry Copy if it fails.
Enter one Sitemap URL. It is trimmed and appended after the blocks without URL validation. The tool does not fetch the URL.
User-agent Choose * (all bots), Googlebot or Bingbot, or select Custom... and type a bot name. A blank custom name writes *; bot names are trimmed. Path rules Add Allow or Disallow and type the path. Paths are trimmed and written without validation; empty paths remain empty rules. A block with no rules writes Disallow:.
User-agent Block
Disallow
robots.txt
User-agent: *
Disallow: /

The starting rule Disallow: / blocks the whole site. Change or remove it before publishing.

Read the full guide robots.txt Explained: Groups, Longest Match and How Crawlers Read the File
Examples, details and FAQ Worked examples, how it compares with other tools, and answers to common questions.

Example: a Typical Site File

Keep the * block, change its rule to Disallow /admin/ and add Allow /admin/public/; add a second block with the custom name GPTBot and Disallow /; enter the sitemap URL. The output is:

User-agent: *
Disallow: /admin/
Allow: /admin/public/

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

Under RFC 9309 the longest matching path wins, whatever the line order, and when an Allow and a Disallow match with the same length the Allow should win. So /admin/public/help may be crawled, /admin/users may not, and GPTBot may fetch nothing. Scrapy’s Protego parser returned exactly that. Python 3.12’s urllib.robotparser did not: it applies the first matching line in file order and reported /admin/public/help as blocked. CPython rewrote the module for RFC 9309 in Python 3.13.14 and 3.14.5 (gh-138907), and Python 3.14.7 allows that URL. If one of your own scripts uses it on an older Python, put the Allow line above the Disallow line; RFC-compliant crawlers read both orders the same way.

Example: Allow One Crawler Only

User-agent: *
Disallow: /

User-agent: Googlebot
Disallow:

Build it with two blocks: the default * block with Disallow /, and a Googlebot block with no rules. A block without rules gets an empty Disallow: line, which allows everything. Googlebot reads only its own group, so the * rule does not apply to it.

Limits and Things robots.txt Does Not Do

  • No validation. Paths are trimmed and otherwise copied as typed. RFC 9309 paths start with /; a value such as admin/ is not a valid rule. Empty paths are written as empty rules.
  • One Sitemap line. For several sitemaps, copy the line and add the others by hand. Crawl-delay and other non-standard lines are not generated; Google ignores Crawl-delay.
  • Wildcards * and $ are part of RFC 9309 and can be typed into a path, for example Disallow: /*?sessionid=; the tool does not check them.
  • Blocking is not removal. A disallowed URL can still appear in search results without a description if other pages link to it. To keep a page out of results, let it be crawled and send noindex, which the meta tag generator can write.
  • Size and errors. RFC 9309 asks crawlers to parse at least the first 500 KiB. If the file returns a 4xx status, crawlers may crawl everything; if it returns 5xx or cannot be reached, they should assume everything is disallowed.

Common robots.txt Patterns

  • Block all bots: User-agent: * + Disallow: /
  • Allow all bots: User-agent: * + Disallow: (empty value)
  • Block a specific directory: Disallow: /admin/
  • Allow Googlebot only: Block *, then add a separate Googlebot block with Disallow:

After generating your file, upload it to your web server root so it is accessible at https://yourdomain.com/robots.txt. Each host and protocol needs its own file: rules at example.com do not cover blog.example.com.

FAQ

What is a robots.txt file?

robots.txt is a plain-text file placed at the root of a website (e.g. https://example.com/robots.txt) that tells web crawlers which pages or directories they are allowed or not allowed to access. It follows the Robots Exclusion Protocol, standardized as RFC 9309.

Does robots.txt block all bots?

No. robots.txt is a convention, not a security measure. Well-behaved crawlers like Googlebot and Bingbot respect it, but malicious bots may ignore it entirely. Use server-level access controls to truly restrict content.

What does 'Disallow: /' mean?

It tells the specified User-agent to not crawl any page on the site. Using 'User-agent: *' combined with 'Disallow: /' blocks all compliant bots from the entire site. This is also the tool's starting state, so change it before you publish the file.

Can I have multiple User-agent blocks?

Yes. Click '+ Add User-agent Block' to add more. A crawler follows only the most specific group that names it, so a Googlebot block replaces the * block for Googlebot instead of adding to it; repeat any shared rules in both blocks.

Where should I put the Sitemap directive?

Anywhere in the file. Sitemap lines are not part of any User-agent group, so their position does not matter; the tool writes the line at the end. It tells crawlers where your XML sitemap is, which helps them discover the pages you want crawled.

Are the paths and URLs I type sent anywhere or saved?

No. The file is built in your browser tab. The tool does not fetch the Sitemap URL, and it does not send your paths or bot names to a server or write them to browser storage, so reloading the page brings back the starting block. When you commit a change (change a field and leave it, pick a bot, add or remove a rule or block), the site's analytics records a usage event with the tool name and the action, not what you typed.