Robots.txt & Sitemap Generator

A robots.txt and an XML sitemap for your site.

Free. Runs in your browser: nothing you enter or open is uploaded.

robots.txt

    

Test a URL against these rules

The tester follows the rules Google publishes: the crawler uses the most specific group that names it, otherwise the * group; the longest matching rule wins, and Allow wins a tie. * matches any characters and $ marks the end of the address. Other crawlers may read edge cases differently. robots.txt asks crawlers to stay out; it does not hide or protect a page.

How to use it

  1. On the robots.txt tab, start from a preset or add Allow and Disallow rules for all crawlers or for named ones, and list your sitemap.
  2. Test a few addresses against the rules, then copy or download the file and upload it to the root of your site, for example example.com/robots.txt.
  3. On the XML sitemap tab, paste your page addresses, add dates if you have them, and download the sitemap. Large lists are split into files of 50,000 addresses with an index.

Questions

Does Disallow stop a page appearing in Google?

Not reliably. Disallow stops crawling, but Google can still list a blocked address it finds through links, without a description. To keep a page out of results, allow crawling and add a noindex robots meta tag instead.

Does Google follow Crawl-delay?

No. Google ignores Crawl-delay. Bing and Yandex read it. If Googlebot is putting too much load on your server, reduce its crawl rate by returning 503 or 429 responses for a short time, as Google's own documentation suggests.

How big can a sitemap be?

Each file may hold up to 50,000 addresses and be up to 50 MB uncompressed. For more, use several sitemap files and a sitemap index that lists them, which this tool creates for you.

Do I need changefreq and priority?

No. Google ignores both. An accurate lastmod date, changed only when the content really changes, is the one optional field Google uses.