You want Google and AI search engines to discover your business. You do not want an automated visitor collecting customer records, internal pricing or unfinished documents. That tension makes AI crawler security a practical website issue. The term supercrawler describes an ambitious automated collection workflow, not a recognised technical security standard. A crawler may follow links, read documents and combine information quickly, but it still depends on what your website makes accessible. Start by deciding which information should be public and which requires controlled access.
Map your public information boundary
Review your website as a visitor without a login. Include pages, linked PDFs, image captions, downloadable spreadsheets, public storage links and structured data embedded in HTML. Marketing teams often approve the visible page while overlooking its attachments. A polished service page can still link to an old presentation containing customer names or a spreadsheet with an internal worksheet. Treat every downloadable file as part of the publication decision.
Create a register with the URL, content owner, purpose and intended audience. For each item, choose one of three outcomes: keep public, replace with an approved public version, or move behind authentication. The register creates a shared decision between sales, operations and whoever maintains the website. It also gives future updates a reference, so a removed sensitive document is not uploaded again from an old folder.
Understand what robots.txt actually controls
A robots.txt file communicates crawling preferences to automated clients that follow the protocol. It cannot authenticate a visitor or prevent a client from requesting a file. RFC 9309 explicitly separates these rules from access authorisation. Use robots.txt to manage cooperative crawling, while placing confidential information behind controls that the server actually enforces.
A crawler name is not proof of identity. A client can send a misleading user-agent string, so avoid treating that string as permission to enter a private area. Review any bot rules with your hosting or security team. The objective is to keep useful public pages available while restricting private resources consistently, including direct file requests and alternative URLs.
Keep Google indexing separate from confidentiality
Google explains that a blocked URL can still appear in search results when discovered elsewhere. A robots.txt restriction therefore does not establish confidentiality. For public material that should disappear from search, use an appropriate indexing rule. For private material, require authentication and authorisation regardless of whether a search engine has ever seen the URL.
Google must be able to read a page to process its noindex directive. Blocking the crawl can prevent it from seeing that instruction. Review crawl access, indexing instructions and server access as separate settings. Do not assume an unlisted page is private simply because it is missing from navigation. Anyone holding a working link may still request it.
Set a deliberate policy for AI discovery
Some operators distinguish search discovery from model training. OpenAI documents separate crawler settings for OAI-SearchBot and GPTBot. This allows different choices for those two purposes. Other services have their own policies and controls, so review each operator's documentation rather than copying a universal block list. Such preferences complement access protection; they do not replace it.
For the content you deliberately publish, clear explanations, stable URLs, descriptive titles and readable HTML support discovery. Our GEO service addresses visibility in AI search alongside conventional search. Build that visibility around approved product facts, useful answers and public evidence. Exclude customer records, passwords and confidential working documents from the public material used to support your marketing.
Check downloads, previews and public API responses
Test download URLs directly in a clean browser session. Check whether a preview link or expired sharing link still exposes content. Review the information returned by public forms and APIs without submitting real customer data. A secure page layout cannot compensate for a response that includes an entire customer object when the visitor only needs a confirmation message.
Give uploaded files an owner and a review date. Before publishing a replacement, inspect its actual contents, including notes, comments and embedded metadata. Use a public export where appropriate rather than exposing an internal original. Remove obsolete copies from the public server as well as from visible links. If sensitive material was already accessible, assess the exposure and remediate the affected data, not just its menu entry.
Maintain the boundary after each update
Make publication checks part of ordinary website maintenance. Compare the upload list with the approved public register, review changed links and verify that private routes still require the expected access. Watch server logs for unusual collection patterns, but interpret them alongside real content and permissions. A high request count can flag investigation; it does not prove that confidential information was exposed.
Discuss new file types with the teams that prepare them. Sales and administration often understand which details belong in a public version better than technical maintainers. Record the approved version and its replacement date. When changing providers, hand over these publication decisions so old exports, presentations and test pages are not restored without review. Assign clear responsibility for both content approval and uploading.
For connected assistants, also read our AI agent security checklist; an agent's access to internal systems creates different questions from public crawling. Discuss a focused review through our AI security service, and follow Blackcarrot Tech on Facebook for practical guidance. Sustainable search visibility works best when your organisation knows exactly which information it has deliberately chosen to publish, update and maintain.