robots.txt, Sitemap and security.txt

Three small files at the root of a site tell machines how to treat it: robots.txt asks crawlers where not to go, sitemap.xml lists the pages worth indexing, and security.txt says who to contact about a vulnerability. Each is easy to get subtly wrong.

robots.txt is a request, not access control

A Disallow line asks well-behaved crawlers to skip a path, and it also publishes that path to anyone who reads the file. Never list a secret location there, and never rely on it to protect anything: use authentication. A single line, Disallow: / under User-agent: *, blocks the whole site from search engines, which is right for staging and a disaster if it reaches production.

User-agent: *
Disallow: /admin/
Allow: /admin/public/

Sitemap: https://example.com/sitemap.xml

Sitemaps and their limits

A sitemap lists canonical URLs you want indexed, one url entry each. A single file may hold at most 50,000 URLs and 50 MB uncompressed; larger sites use a sitemap index. Only add a lastmod date if it is accurate for every URL, because search engines learn to ignore a date that is always wrong.

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/about</loc>
    <lastmod>2026-10-04</lastmod>
  </url>
</urlset>

security.txt (RFC 9116)

Serve /.well-known/security.txt over HTTPS so a researcher who finds a vulnerability knows where to report it. Contact and Expires are required. Keep Expires under a year away and renew it on a schedule, because a file past its expiry is treated as abandoned.

Contact: mailto:security@example.com
Expires: 2027-06-30T23:59:59.000Z
Canonical: https://example.com/.well-known/security.txt

Opting out of AI training crawlers

Several AI companies publish a user-agent token for the crawler that collects training data, such as GPTBot, ClaudeBot and Google-Extended, and say they honour a Disallow for it. This is a request: crawlers that ignore robots.txt are unaffected, and blocking a token does not remove content already collected.

Open the robots.txt, sitemap & security.txt generator

Frequently asked questions

Does Disallow in robots.txt remove a page from Google?

Not reliably. It stops crawling, but a blocked URL can still appear in results if other sites link to it. To keep a page out of the index, let it be crawled and serve a noindex directive.

Where must these files be placed?

robots.txt and sitemap.xml at the root of the host they describe, and security.txt at /.well-known/security.txt. Each host and subdomain needs its own robots.txt.

Do I need a sitemap for a small site?

Not strictly, since crawlers follow links. It still helps new or poorly linked pages get found sooner, and it is cheap to provide.

Guides