Configuring robots.txt comes down to a short sequence: place the file, learn the directives, write your rules, decide on AI crawlers and test the result. The file is small, yet one wrong line can stop search engines from crawling your entire site or expose content you meant to keep private.
Step 1: Put the file at your domain root
When a search engine bot reaches your site, it looks for robots.txt at the root of the domain before anything else. The file lists which bots may visit which parts of the site, and well-behaved crawlers follow it, although compliance is voluntary.
Step 2: Learn the four core directives
- User-agent: names the crawler a group of rules applies to.
- Disallow: blocks a specific path.
- Allow: opens a specific path inside a directory you have disallowed.
- Sitemap: points crawlers to your XML sitemap.
These four cover nearly everything a typical business website needs.
Step 3: Block the areas that should stay out
Most business sites follow the same pattern. Block admin areas and login pages, internal search result pages, duplicate content variations and any development or staging directories, while leaving every important content page open and listing your sitemap location.
Step 4: Decide how to treat AI crawlers
AI crawlers have given robots.txt a bigger role. Many sites now add rules for AI training crawlers such as GPTBot and Google-Extended, and choosing which ones to allow or block has become a real business decision.
Step 5: Test the rules before and after you publish
Check your configuration with the robots.txt tester in Google Search Console, and run specific URLs against your rules to confirm that key pages stay reachable and private pages stay blocked. After any change, watch your crawl stats and indexing reports for surprises.
Step 6: Use noindex where blocking is not enough
Robots.txt rules stop crawling, yet they do not always keep a page out of search results. For pages that must never appear, add a noindex meta tag or an X-Robots-Tag header alongside your robots.txt rules.
Tags
Frequently Asked Questions
Where does the file go in the first step?
Robots.txt is a plain text file that sits at the root of your domain. It tells web crawlers which pages or sections they may or may not visit, and it is the first file a search engine bot reads before it crawls the site.
Will a Disallow rule keep a page out of search results?
Not always. Robots.txt stops crawling but not indexing, so Google may still index a blocked URL that other sites link to, even without reading its content. To keep a page out of the index, pair robots.txt with a noindex meta tag or an X-Robots-Tag header.
What goes wrong if I skip the testing step?
A bad rule can block important pages from crawling and indexing, hide the whole site from search engines, or open private areas to crawlers. Test every change with care and watch your indexing reports afterward.
Enjoyed this article?
Share it with your network









