What robots.txt does

A robots.txt file sits at the root of your site, for example yourdomain.com/robots.txt, and tells crawlers which paths they may crawl. Well-behaved search engines and AI crawlers read it before fetching pages.

The basic syntax

  • User-agent: names the crawler the rules apply to. * means all.
  • Disallow: a path the crawler should not fetch.
  • Allow: an exception inside a disallowed path.
  • Sitemap: the full address of your XML sitemap.

Create one with Serpgy

  1. Open the Robots.txt Generator.
  2. Choose the default rule for all crawlers.
  3. Add any folders to block, such as /admin/ or /cart/.
  4. Decide whether to allow or block specific AI crawlers.
  5. Enter your sitemap address, copy the file and upload it to your site root.

A simple example

User-agent: * / Allow: / / Disallow: /admin/ / Sitemap: https://yourdomain.com/sitemap.xml. This allows everything except the admin area and points crawlers to your sitemap.

What robots.txt does not do

  • It is not security. The file is public and blocked URLs can still be visited by anyone. Protect private areas with authentication.
  • It does not reliably remove pages from search results. A blocked page can still appear if others link to it. Use a noindex tag on a page crawlers are allowed to fetch, or remove the page.
  • Not every bot obeys it. It is a request, not enforcement.

Common mistakes

  • Leaving Disallow: / in place after launching a site, which blocks the whole site.
  • Blocking CSS and JavaScript files that pages need to render properly.
  • Placing the file in a subfolder. It must be at the root.
  • Forgetting that paths are case-sensitive.

AI crawlers

You can address AI-related user agents separately, for example GPTBot or Google-Extended, to allow or disallow them independently of normal search crawling. Check each provider's documentation for the current user-agent names.