Robots.txt & AI-bot directives audit
We check your robots.txt and the directives for AI bots: syntax correctness, whether anything important is accidentally blocked or anything unnecessary is opened, the sitemap reference, and the settings for AI crawlers (GPTBot, Google-Extended, CCBot, etc.) — whether to let them into your content or not. And honestly upfront: robots.txt is a crawl directive, not protection; it does not guard sensitive data, and some bots ignore it.
Robots.txt & AI-bot directives audit — overview

A robots.txt & AI-bot directives audit is a one-time narrow check of the file that tells search and other robots what they may crawl and what they may not: the correctness of syntax and directives (User-agent, Disallow, Allow), whether important sections are blocked by mistake (a common and costly error) and whether unnecessary ones are opened (service, technical pages), whether there is a sitemap reference, whether there are conflicts and contradictory rules, and correctness for different crawlers. Separately — the directives for AI bots: access settings for neural-network crawlers (OpenAI's GPTBot, Google-Extended, Common Crawl's CCBot, PerplexityBot, etc.) — whether you want your content used for training and in AI answers, and whether it is set up correctly. The goal is to put the crawl rules in order for your needs. Honestly about the main point, and this matters: robots.txt is an instruction to robots about which pages to crawl, NOT a means of protection or access control. A page blocked in robots can still get into the index if there are links to it; and sensitive data (user accounts, documents) is not guarded by robots.txt — it must be protected with a password and server settings, not an entry in this file. And honestly: robots.txt is an agreement, not a server-level ban: well-behaved bots (Yandex, Google, GPTBot, etc.) obey it, but malicious or some third-party scrapers may ignore it. On AI: a robots ban does not remove your content from what AI is already trained on — it only affects further collection, and not all AI crawlers respect such directives equally. It is a snapshot in time: new bots and directives appear. For a full check, access to the current robots.txt at the site root is needed (and for staging, multisite or subdomains — to their variants); without access we check the publicly available version. What you get: a report with errors and risks in robots.txt and the AI directives by priority and recommendations (what to block, what to open, how to set it up for AI); implementation is separate small work, usually on the developer side, and the report by itself changes nothing without implementing the rules. If the site is simple and robots is standard, we will honestly say a basic check is enough. Picture this: instead of "robots seems to be there" you learn that a single Disallow line accidentally blocks the whole catalog section, there is no sitemap reference, and GPTBot and Google-Extended are not configured — meaning the question of AI using your content is simply left to chance. The base price starts from 6,000 ₽.
Problems we solve
- You are not sure robots.txt does not block too much or open something important.
- Part of the site disappeared from search — you suspect robots.
- You want to control whether AI uses your content (GPTBot, etc.).
- A site migration or redesign — you need to check the crawl rules.
What's included in the Robots.txt & AI-bot directives audit service
- Syntax and correctness of directives (User-agent, Disallow, Allow)
- Accidentally blocked important sections and opened unnecessary/service ones
- Sitemap reference and consistency with it
- Conflicts and contradictory rules for different robots
- Directives for AI bots (GPTBot, Google-Extended, CCBot, PerplexityBot, etc.)
- A check that robots is not used as "protection" for sensitive data
- Behavior for the main search and AI crawlers
- A prioritized report with errors and recommendations
What you get
- Clear understanding of what is misconfigured in robots.txt
- A list of errors: what is blocked needlessly, what is opened unnecessarily
- A decision on AI bots: whether to allow them and how to specify it
- Understanding that robots is about crawling, not data protection
How the work goes: steps
- We clarify what should be blocked or open and whether to allow AI bots
- We check the syntax, conflicts, important sections and AI directives
- We prepare a report with errors and recommendations and review it with you
Why PDV Expert
- Fixed price and timeline — no surprises on the invoice.
- Report and recommendations in plain language — clear without a technical background.
- In touch at every step and answering questions about the result.
FAQ
Will robots.txt protect blocked pages from outsiders?
No. robots.txt only asks robots not to crawl pages, but it is not protection: a blocked page can still get into the index via links, and malicious bots ignore the file. User accounts and documents must be protected with a password and server settings, not robots.
If I block GPTBot, will my content disappear from ChatGPT?
Not entirely. A robots ban only affects further data collection but does not remove what AI is already trained on. Also, not all AI crawlers respect such directives equally. We will honestly show what the setting actually achieves and what it does not.
Is this the same as an indexation or sitemap audit?
No. robots.txt is about crawl rules (what robots may do), the sitemap is about the map of needed pages, and an indexation audit is about what is actually in the index. They are related, but they are different checks; here the focus is on robots.txt and AI directives.
About the provider
The «Robots.txt & AI-bot directives audit» service is provided by PDV Expert — a team specialising in «Diagnostics & monitoring». We work under contract and deliver a written report with recommendations.