Context Theory Get your growth audit

Answer

Should you block AI crawlers from your site?

For a business that wants to be found, almost certainly not. Blocking removes you from answers without removing you from competition.

Not if you sell services and want enquiries. Blocking removes you from generated answers without removing the question, so a competitor is cited instead. It is a reasonable choice only where the content itself is the product.

The instinct to block is understandable and the reasoning usually transfers badly from a different business model. Publishers whose product is the content itself have a genuine grievance: an answer that reproduces their reporting removes the visit that funded it, and blocking is a rational defence of the thing they sell. A service business is in the opposite position. Its content is not the product; it is the advertisement for the product, and the objection to it being reproduced is much weaker.

What blocking achieves for such a business is narrow and mostly unwanted. The questions people ask about your industry, your area and your kind of work continue to be asked and continue to be answered. Blocking removes you from the pool of possible sources for those answers. It does not remove the answer, and it does not remove your competitors from it. You have withdrawn from a contest that carries on without you.

Against that sits a real concern worth taking seriously rather than dismissing: crawling costs bandwidth, some crawlers behave badly, and a business may reasonably object on principle to its material being used to train systems it did not consent to. Those are legitimate and they are separable. Crawl load is addressable with rate limiting rather than exclusion. Training and answering are frequently distinguishable by crawler, since several operators publish separate agents for each and honour rules that apply to one and not the other — which means the position of allowing retrieval while refusing training is often expressible rather than hypothetical.

That distinction is where a considered policy usually lands, and it needs checking rather than assuming, because the agent names and the operators' stated behaviour change. The durable version of the position is: permit crawlers that fetch in order to answer questions and cite sources, restrict those whose stated purpose is bulk collection for training, and rate-limit anything that misbehaves. Implementing it requires reading the current documentation of the operators that matter rather than copying a blocklist from a forum post, because a copied list goes stale in both directions.

There is a partial-blocking mistake worth avoiding. Businesses sometimes allow crawling of a homepage and block the sections that contain the actual substance — pricing, service detail, case material — which produces the worst available outcome: the site is present but has nothing extractable, so it is passed over in favour of a competitor who published detail. If the decision is to participate, the pages carrying the facts are exactly the ones that must be reachable.

The honest summary is that this is a business-model question rather than a technical one. If people pay you for access to what you publish, blocking is worth serious consideration. If what you publish exists to bring you enquiries, blocking removes a distribution channel and protects nothing you were selling.

Blocking a crawler does not stop the question being asked; it only guarantees that the answer is assembled from your competitors.

Answer Production Engine, Context Theory

Related questions

Do these crawlers actually obey robots rules?

Major operators publish their agent names and state that they honour standard exclusion rules, and compliance in practice is generally reported as good for the named ones and unreliable at the edges. Which means exclusion is a request rather than a control. If genuinely preventing access matters, that is an access-control problem — authentication or blocking at the server — not a robots file question.

Does allowing crawling mean our content trains a model?

Not necessarily, and the answer depends on the operator and the agent. Several publish distinct agents for retrieval and for training-data collection precisely so that the two can be permitted separately. Establishing what applies to your site means reading the current documentation for the operators you care about, since both the agent names and the stated behaviours have changed more than once.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
AI-cited sources that also rank in the Google organic top 1010%Category-wide
Same, when an AI Overview is present83%Category-wide
Organic CTR with an AI Overview vs without0.61% vs 1.62%Category-wide

2026 generative engine citation study · fewer than · verified

2026 zero-click search analysis · verified

2026 SERP analysis · derived: 1.62 → 0.61 = −62.3% · verified

What is specific to this page.

Evidence
Kind Claim Check it against
SoftwareBlocking removes a site from the pool of sources for questions that continue to be asked and answered, so it withdraws the business from a contest that proceeds using its competitors as sources instead.Asking a generated assistant a question in the business's category and observing which competitors are cited.
ConstraintSeveral operators publish separate crawler agents for retrieval and for bulk training-data collection and state that exclusion rules are honoured per agent, so permitting one while refusing the other is expressible rather than hypothetical.The current published crawler documentation of the operators concerned, which lists agent names and stated purposes.
SoftwareAllowing a homepage while blocking the sections carrying pricing, service detail and case material produces a site that is present but has nothing extractable, which is worse for the business than either full participation or full exclusion.The site's own exclusion rules, checked against which pages carry its checkable facts.
ConstraintExclusion rules are a request rather than an access control, so a business that genuinely needs to prevent retrieval must use authentication or server-level blocking rather than a robots file.Server logs for requests from named crawler agents to paths the exclusion rules disallow.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one