AIboosting a SiteBoosting brand
Back to explainers AI findability

Robots.txt: decide who may go where on your site

By Stefan Hendriks

Robots.txt is a small file that tells crawlers which parts of your site they may and may not visit. Set incorrectly, you accidentally block your own findability. Read on.

Every serious website has an invisible little file that plays an important role: robots.txt. It is the doorkeeper of your website. It tells the crawlers of search engines and AI which parts they may visit and which they may not. Small file, big consequences if it is set wrong.

What is robots.txt?

Robots.txt is a text file in the root of your website, reachable at yourdomain.com/robots.txt. It contains instructions for crawlers: which folders or pages they may visit and which they should skip. It is a voluntary standard, but the large, trustworthy crawlers abide by it.

Where does it often go wrong?

A classic mistake is that during a website’s build everything is blocked for crawlers, and after going live this is forgotten and never reset. The result: the site appears nowhere. Another mistake is accidentally blocking important folders so crawlers cannot read the content.

Robots.txt in the AI age

Robots.txt is more relevant than ever, because AI crawlers such as GPTBot, ClaudeBot and PerplexityBot also abide by it. With it you can decide whether AI systems may use your content. For most businesses that want to be found by AI, this means: allowing these crawlers in.

What does this mean for you?

Robots.txt is a powerful but sensitive instrument. A single wrong rule can block your entire findability. A correct setting gives you control over who may read your content, including the growing group of AI crawlers.

Need help with your AI findability? To SiteBoosting