Know which crawler does what

AI companies publish the names of their crawlers. OpenAI, for example, documents separate agents for training data, for search features and for fetching pages when a user asks. Blocking the training crawler does not have to block the search crawler. Google-Extended controls whether your content is used for some Google AI models, separately from Googlebot for search.

A sensible default

If your site is marketing content meant to win customers, allow the search and user-triggered crawlers so you can be cited. Decide on training crawlers based on how you feel about your content being used to train models. Protect genuinely valuable material, such as paid resources or client data, behind logins rather than relying on robots.txt.