Three small files at the root of a site tell machines how to treat it: robots.txt asks crawlers where not to go, sitemap.xml lists the pages worth indexing, and security.txt says who to contact about a vulnerability. Each is easy to get subtly wrong.
A Disallow line asks well-behaved crawlers to skip a path, and it also publishes that path to anyone who reads the file. Never list a secret location there, and never rely on it to protect anything: use authentication. A single line, Disallow: / under User-agent: *, blocks the whole site from search engines, which is right for staging and a disaster if it reaches production.
User-agent: *
Disallow: /admin/
Allow: /admin/public/
Sitemap: https://example.com/sitemap.xmlA sitemap lists canonical URLs you want indexed, one url entry each. A single file may hold at most 50,000 URLs and 50 MB uncompressed; larger sites use a sitemap index. Only add a lastmod date if it is accurate for every URL, because search engines learn to ignore a date that is always wrong.
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/about</loc>
<lastmod>2026-10-04</lastmod>
</url>
</urlset>Serve /.well-known/security.txt over HTTPS so a researcher who finds a vulnerability knows where to report it. Contact and Expires are required. Keep Expires under a year away and renew it on a schedule, because a file past its expiry is treated as abandoned.
Contact: mailto:security@example.com
Expires: 2027-06-30T23:59:59.000Z
Canonical: https://example.com/.well-known/security.txtSeveral AI companies publish a user-agent token for the crawler that collects training data, such as GPTBot, ClaudeBot and Google-Extended, and say they honour a Disallow for it. This is a request: crawlers that ignore robots.txt are unaffected, and blocking a token does not remove content already collected.
Open the robots.txt, sitemap & security.txt generator
Not reliably. It stops crawling, but a blocked URL can still appear in results if other sites link to it. To keep a page out of the index, let it be crawled and serve a noindex directive.
robots.txt and sitemap.xml at the root of the host they describe, and security.txt at /.well-known/security.txt. Each host and subdomain needs its own robots.txt.
Not strictly, since crawlers follow links. It still helps new or poorly linked pages get found sooner, and it is cheap to provide.