What robots.txt does
A robots.txt file sits at the root of your site, for example yourdomain.com/robots.txt, and tells crawlers which paths they may crawl. Well-behaved search engines and AI crawlers read it before fetching pages.
The basic syntax
- User-agent: names the crawler the rules apply to. * means all.
- Disallow: a path the crawler should not fetch.
- Allow: an exception inside a disallowed path.
- Sitemap: the full address of your XML sitemap.
Create one with Serpgy
- Open the Robots.txt Generator.
- Choose the default rule for all crawlers.
- Add any folders to block, such as /admin/ or /cart/.
- Decide whether to allow or block specific AI crawlers.
- Enter your sitemap address, copy the file and upload it to your site root.
A simple example
User-agent: * / Allow: / / Disallow: /admin/ / Sitemap: https://yourdomain.com/sitemap.xml. This allows everything except the admin area and points crawlers to your sitemap.
What robots.txt does not do
- It is not security. The file is public and blocked URLs can still be visited by anyone. Protect private areas with authentication.
- It does not reliably remove pages from search results. A blocked page can still appear if others link to it. Use a noindex tag on a page crawlers are allowed to fetch, or remove the page.
- Not every bot obeys it. It is a request, not enforcement.
Common mistakes
- Leaving Disallow: / in place after launching a site, which blocks the whole site.
- Blocking CSS and JavaScript files that pages need to render properly.
- Placing the file in a subfolder. It must be at the root.
- Forgetting that paths are case-sensitive.
AI crawlers
You can address AI-related user agents separately, for example GPTBot or Google-Extended, to allow or disallow them independently of normal search crawling. Check each provider's documentation for the current user-agent names.