# How Website Owners Can Protect Their Content From AI Training Without Losing Search Visibility
For years, website operators have been caught in a frustrating dilemma: allow search engines and AI companies to freely crawl and use your content for machine learning, or block them entirely and risk losing your place in search results. The root of this problem lies in the fact that many of the web’s biggest crawlers serve a dual purpose — they build search indexes **and** scrape data to train artificial intelligence models at the same time. Choosing to opt out of one meant opting out of both.
That landscape is shifting. A major infrastructure provider has introduced a new, granular setting that lets website owners keep their pages indexed by search engines while explicitly declining to let those same crawlers use their content for AI model training. The update has already gained support from several of the internet’s most influential organizations, including Apple, Google, and Microsoft.
## The Core Problem: Shared Crawlers Create a False Choice
When a single crawler handles both search indexing and AI training, website owners are forced into an all-or-nothing decision. Many people assume that search and AI training are interchangeable — if you block one, the other goes away too. In practice, that has meant that site owners wanting to protect their content from being used to build AI models had no clean way to preserve their organic search traffic.
This issue matters differently depending on how a website operates. Some sites rely entirely on advertising revenue, which depends on actual human visits. Others operate on subscriptions or direct customer relationships. In all cases, the fundamental challenge is the same: crawlers used for training bypass the site’s economic model by consuming content without generating meaningful traffic in return.
## The New Disallow AI Training Setting
The newly introduced setting, called **Disallow AI Training**, works at the network level rather than relying solely on voluntary compliance. Here is how it functions:
– It publishes a **no-training preference** directly into the site’s `robots.txt` file.
– Crawlers that have agreed to respect this setting are allowed to continue indexing the site for search purposes but are prevented from using the content for training or fine-tuning AI models.
– Training-only crawlers — those that do not contribute to search indexing — are blocked entirely.
This approach is fundamentally different from simply blocking a crawler with a `Disallow` rule, which would remove both search and training access simultaneously. The new setting preserves discoverability while drawing a clear line around AI training.
## Why Robots.txt Alone Was Not Enough
A `robots.txt` directive is easy for anyone to publish, but it has significant limitations. It cannot reliably identify which bot is visiting, determine the purpose of the crawl, or enforce compliance when a crawler chooses to ignore the instructions. The new system works because the infrastructure provider can verify the identity of each crawler, classify its behavior, and take action against those that disregard the stated preferences. It also tracks and reports how each operator responds to these signals.
## Three Types of Crawler Behavior
The updated system classifies all crawler activity into three distinct behaviors, each of which can be controlled independently:
1. **Search** — crawling activity aimed at building or updating a search index.
2. **Training** — crawling activity intended to train or fine-tune an AI model.
3. **Agent** — user-directed agents that visit pages on behalf of a human, such as chat-based fetch bots or browser-use automation tools.
Mixed-use crawlers perform both Search and Training simultaneously. By separating these behaviors into distinct controls, website owners now have the flexibility to allow one while restricting the other.
## Updated Settings and What They Mean
With the introduction of Disallow AI Training, the available configuration options have expanded:
– **Allow** — Every crawler is permitted unless restricted by another rule or firewall policy.
– **Disallow AI Training** — Publishes the no-training preference in `robots.txt`. Search-capable mixed-use crawlers that honor the standard remain active for indexing. All other training crawlers, including training-only crawlers from major AI companies, are blocked. This setting does not affect search.
– **Block on pages with ads** — All crawlers, including mixed-use crawlers, are restricted only on pages detected as serving advertisements.
– **Block** — All crawlers are denied access, including those used for search indexing.
It is worth noting that Disallow AI Training only applies to the Training behavior, not Search or Agent. The reason is that agents do not create the same search-discoverability tradeoff, and the web does not yet have a universally supported standard for expressing agent-level preferences.
## What Changes on the Rollout Date
Several important adjustments take effect when the new system launches:
– The existing **Block** and **Block on pages with ads** settings now apply to mixed-use crawlers as well, meaning they impact both search and training access.
– A legacy **”Block AI Bots”** option is being phased out in favor of the more detailed Search, Training, and Agent controls.
– The previous Managed Robots.txt system is being replaced by a new Bot Preference Sync mechanism.
– New domains being added will receive one of two recommended configurations based on whether the site uses advertising revenue.
– Current customers will see their existing preferences automatically migrated to the new system.
In most cases, website owners do not need to take any action — their current settings carry forward seamlessly.
## Migration Tables for Existing Customers
For domains that previously used the simplified “Block AI Bots” controls, the migration maps directly to the new settings:
| Legacy Setting | Search Behavior | Training Behavior | Agent Behavior |
|—|—|—|—|
| Not enabled (default) | Allow | Allow | Allow |
| Block | Allow | Disallow AI Training | Block on pages with ads |
| Block on pages with ads | Allow | Disallow AI Training | Block on pages with ads |
For domains that already used the granular Search, Training, and Agent controls, the practical effect of prior decisions is preserved under the new definitions. Previous Training selections of Block or Block on pages with ads transition to Disallow AI Training.
## Recommended Settings for New Domains
When onboarding a new domain, customers are presented with a choice between two preset configurations:
| Setting | Non-Ad-Supported Site | Ad-Supported Site |
|—|—|—|
| Preference Sync | Enabled | Enabled |
| Search | Allow | Allow |
| Training | Allow | Disallow AI Training |
| Agent | Allow | Block on pages with ads |
The reasoning is straightforward: ad revenue depends on real human traffic. AI training replaces a visit with a generated answer, and agents fetch pages without a human present to see advertisements. For this reason, sites that monetize through ads receive more protective default settings.
## Which Crawlers Honor the New Standard
Several major search and technology companies have either already implemented or committed to honoring the Disallow AI Training setting:
– **Applebot** allows site owners to opt out of training via a `Disallow` rule for “Applebot-Extended” in `robots.txt`. The company has confirmed that disabling training does not affect search rankings. Apple currently provides limited URL-level transparency tools but has indicated that more advanced features are under development.
– **Googlebot** supports training opt-outs through a `Disallow` rule for “Google-Extended” and offers a toggle in its webmaster portal to exclude content from AI-generated search results. Google provides detailed reporting on both search and AI summary performance and has confirmed that blocking Google-Extended does not impact search ranking.
– **Bingbot** offers controls through its webmaster portal and currently supports AI training preferences via the `NOARCHIVE` meta tag. Microsoft is working on adding full robots.txt compliance for the no-training preference, with a targeted rollout in early 2027. The company has confirmed that using `NOARCHIVE` does not affect search ranking.
Additionally, several AI-focused organizations — including Amazon, Anthropic, Meta, and OpenAI — operate separate crawlers for search and training. Because these functions are already split across different user-agent identifiers, it is possible to block the training crawler without affecting search indexing.
## The Road Ahead: AI Summaries
While the new training controls address an important concern, they are not the final piece of the puzzle. AI summaries represent a separate and distinct challenge. Unlike training, which affects model development behind the scenes, summaries directly shape how audiences discover, evaluate, and interact with a website’s content.
The operators identified as Accountable either already provide or are actively building the tools needed to opt out of AI summaries. However, a site-wide on-or-off switch for summaries is still a blunt instrument. Different businesses have different goals — a publisher may prioritize audience size, while a retailer may prefer fewer but higher-intent visitors. The next phase of development aims to give site owners more nuanced control over exactly how much of their content appears in summaries and under what circumstances.
Industry data suggests that summaries are reshaping user behavior in complex ways. More than half of consumers now read AI-generated summaries within search results, and those users are significantly more likely to end their search session without clicking through to any website. At the same time, consumers referred by AI-powered search convert at rates three to five times higher than those arriving through traditional search results. The balance between visibility and engagement is shifting, and having the right controls will be essential for every website operator.
## The Accountable Designation
To help website owners understand which crawler operators respect their choices, the infrastructure provider introduced an **Accountable** designation. To qualify, a bot operator must meet or commit to meeting four key requirements:
1. A functioning mechanism for site owners to opt out of AI training, whether through `robots.txt` or an equivalent standard.
2. A dedicated mechanism for opting out of AI summaries, both directly with the operator and through the platform in the near future.
3. URL-level visibility showing which pages were made available for training, alongside metrics demonstrating how content performs in search results.
4. A clear assurance that opting out of AI training will not degrade traditional search visibility or ranking.
The Accountable label is designed to provide transparency and trust, giving site owners confidence that their preferences will be respected.
## Frequently Asked Questions
**Q: Do I need to change my existing settings?**
A: In most cases, no. Your current preferences are automatically migrated to the new system without any action required on your part.
**Q: Will blocking AI training affect my search rankings?**
A: No. The companies that have committed to this standard — Apple, Google, and Microsoft — have all confirmed that opting out of AI training does not impact traditional search results.
**Q: What happens if a crawler ignores my preferences?**
A: The system detects non-compliant crawlers and blocks them. The provider also publicly tracks and reports on how each operator responds to these signals.
**Q: Can I still block all crawlers entirely?**
A: Yes. The existing **Block** setting remains available and will prevent all crawlers, including those used for search, from accessing your site.
**Q: Will there be a setting to control AI agents in the future?**
A: Not immediately. The team plans to revisit agent controls once web standards such as `ai-prefs` mature and gain broader adoption.
**Q: Are these controls available on all plans?**
A: Yes. The new settings are available to all customers regardless of their plan level and can be configured at the domain level in the security settings dashboard.
**Q: How does the system handle pages that serve ads?**
A: There is no separate Disallow AI Training option for ad pages because ad-serving pages change too frequently and in too large a volume to enumerate in `robots.txt`. Instead, the system uses detection-based blocking for the “Block on pages with ads” setting.
## Conclusion
The introduction of granular AI training controls represents a meaningful step toward restoring balance between website operators and the automated systems that crawl their content. By allowing site owners to protect their material from AI training without sacrificing search visibility, the new framework addresses a tension that has frustrated publishers and businesses for years.
This progress depends on collaboration between infrastructure providers, content creators, technology companies, and standards organizations working together to build open, interoperable systems. The Accountable designation, the new Disallow AI Training setting, and the evolving tools around AI summaries all point toward a future where the people who create the web have genuine agency over how their work is used.
Website owners can access these controls today through their domain’s security settings, and the system is designed to work seamlessly regardless of whether they are running a personal blog or a large commercial site.
Thank you for reading



