Free Tool

Robots.txt Generator

Build a valid robots.txt file with custom rules, your sitemap, and explicit AI bot controls — ready to upload to your domain root.

Rules for all crawlers (*)
AI crawlers
Your robots.txt
Set your rules on the left, then click "Generate robots.txt."
Upload the output as robots.txt to your domain root. Try the Robots.txt Validator next.

What robots.txt actually controls

robots.txt is a simple, plain-text instruction file at the root of your domain that tells web crawlers which paths they're allowed or not allowed to request. It's a voluntary standard — well-behaved crawlers from Google, Bing, and the major AI labs respect it, but it isn't a technical access-control mechanism. A private page still needs real authentication, not just a Disallow rule, to actually stay private.

Getting it wrong is a common, high-impact mistake: a single misplaced Disallow: / under the wrong user-agent group can accidentally block your entire site from Google, often for months before anyone notices the traffic drop.

Why AI bot rules deserve their own decision

Beyond the classic Googlebot/Bingbot question, robots.txt now controls a second, separate decision: whether AI crawlers like GPTBot, ClaudeBot, and Google-Extended can access your content for model training, and whether answer-engine crawlers like PerplexityBot and ChatGPT-User can access it for live citation. These are genuinely different tradeoffs — allowing training crawlers has different implications than allowing citation-oriented ones — which is why this generator lets you set them independently instead of treating "AI bots" as one blanket category.

How to use this robots.txt generator

Frequently asked questions

What is a robots.txt file for?

robots.txt is a plain text file at your domain's root that tells web crawlers which parts of your site they're allowed or not allowed to access. It's a voluntary standard that well-behaved crawlers, including Googlebot, respect.

Does blocking a page in robots.txt remove it from Google's index?

Not necessarily. Blocking crawling prevents Google from reading the page's content, but if other pages link to it, Google can still index the URL with limited information. To fully prevent indexing, use a noindex meta tag on the page itself, which requires the page to be crawlable in the first place.

Where do I upload my robots.txt file?

At the root of your domain, so it's reachable at yourdomain.com/robots.txt exactly. It must be at the root, not in a subfolder, for search engines to find and respect it.

Should I include AI bot rules in my robots.txt?

It's worth deciding deliberately rather than leaving it to chance. If you want your content citable by ChatGPT or Perplexity, allow their crawlers explicitly; if you want to prevent your content being used for AI model training, add specific disallow rules for bots like GPTBot and Google-Extended.

Start building your content library in under 8 minutes.

Your competitors are publishing every week. Every week you don't is a week of organic traffic going to them.

No Credit Card Required.