A website owner who lets a search engine find an article has made one decision. Whether that article should help train a commercial AI model is another. Whether an assistant should retrieve it to answer a reader's question is another still. The requests may reach the same page, but the owner can have different reasons for accepting each one.
The open web needs to make those distinctions easier to express and harder to ignore. Putting work online for people to read should not count as consent for every business that can find a use for it. At the same time, permission cannot become an unlimited power to prevent other people from learning, quoting, or doing research.
A site owner needs to be able to refuse a specific use and find out whether the refusal was respected. The new controls arriving on the web deserve to be judged on how well they allow both.
A public page has several possible uses
Consider the difference between putting an article in a search index and using it to train a model. The index helps a service retrieve information about the page. Training changes a model through exposure to material. Fetching the page during a chat is a third use, even when that answer appears in the same window as answers made without a fresh visit to a page.
These differences matter to the person paying to publish. An owner might welcome discovery, permit some summaries, and decline training. Another might make the opposite bargain. A publication with freely licensed material could welcome reuse while objecting to a crawler that overwhelms its server. None of those positions requires treating AI as a single visitor with a single purpose.
Some AI companies already draw these lines. OpenAI documents separate roles:
- GPTBot is for potential training use.
- OAI-SearchBot is for search.
- ChatGPT-User handles certain user-directed actions, and OpenAI says robots.txt rules may not apply to those visits.
Blocking one named bot therefore leaves other access questions unresolved, even within the same company's services.
A site owner should not have to learn each crawler's name to understand what they have agreed to. Even a simple settings page should make the separate uses clear. One large switch labeled AI can conceal as much as it clarifies.
The Authors Guild offers a useful example from contracts. Its AI model clauses, updated in April 2026, separate training rights from uses such as retrieval-based summaries. The proposed terms give authors a way to discuss each use before agreeing to it. Existing contracts may treat those rights differently.
There is another person to consider, too. A website operator may possess the files without holding every right in the work. A publisher's agreement with a freelancer can matter as much as the publisher's agreement with an AI company. Moving the decision into a dashboard does not settle who is entitled to make it.
When a preference becomes a real choice
On September 15, Cloudflare announced its Disallow AI Training setting and an Accountable designation for operators of mixed-use crawlers. The idea is to let a site remain discoverable while refusing training. The designation covers tools already in use and promises of more to come, including ways to control and track how content is used.
Site owners could gain more control from those promises, but they still need to know which tools they can use today and which ones they must wait for.
Google now lets sites opt out of generative Search features without affecting normal Search ranking or inclusion, and has separate controls for training. The claim that publishers everywhere still face one indivisible choice between search and AI is already too crude.
Yet adding separate settings is only a start. The person selecting one needs to understand its scope. Does it govern a future crawl, a use of material already collected, or the display of material in an answer? An owner can agree to a label that blurs those uses and still have little idea what the choice means.
The familiar robots.txt file illustrates the problem. The Robots Exclusion Protocol says its rules are not access authorization. The file tells cooperating crawlers what an owner wants them to do. It does not check who each visitor is or stop a request from reaching the server.
For that request to have an effect, the crawler operator has to honor it. An operator offering an opt-out should explain which systems it covers, when it takes effect, and what evidence a publisher can inspect. A server log can help show that a bot visited. It cannot, by itself, prove what happened inside a training process afterward.
The burden should not fall entirely on the person with the smallest technical staff. A right that takes a specialist to exercise is less useful to the independent writer maintaining a site after work than to a company with a legal department. The same applies to finding out whether the right was respected. A polished settings page cannot replace a clear record and a way to dispute it.
Can the owner change their mind later? Future access can be reconsidered. Material already used to train a model presents a different problem, and a provider should explain the difference. Selling an access switch as a rewind button would be another failure of consent.
Permission still has limits
The strongest objection is that a web organized around permission could become a place where only large companies can afford to read at scale. Lisa Macpherson of Public Knowledge has warned that broad opt-in requirements and licensing costs can damage fair use, restrict access, and reinforce the position of powerful intermediaries.
She is right to make that objection. A small research project cannot negotiate with every website it studies. Nor should a publisher gain a veto over criticism simply because criticism involves reading and quoting its work. Some uses are permitted without the owner's agreement. In the United States, fair use involves factors including purpose, the nature and amount of the material, and market effects. A preference file cannot replace that analysis.
The reverse shortcut is no better. A page's availability does not answer all those questions either. Being able to fetch a file tells us that a server returned it. It does not tell us that the author endorsed every later use, or that every use meets the legal test.
I would rather see a narrower, workable consent system than a grand promise of control that cannot survive contact with ordinary reading. Start with services making clear commitments about their own conduct. An owner's expressed wishes deserve respect within the limits the law allows. Give researchers and the public a place in the discussion before the commercial agreements harden around them.
Payment belongs in that discussion, but it cannot do all the work. A person can agree to free reuse. They can decline a paid offer. They can accept one use and refuse another. Reducing consent to a royalty rate would erase the very choice the rate is supposed to recognize.
The next improvement should be small enough to describe without a sales presentation. A site owner refuses a named use. Readers can still find the page through search, where that was promised. The operator records what the refusal covers, and there is a way to check what happened. The page can remain open, and the answer can still be no.



