What is a mixed-use crawler?
.png)
Why one company runs one fleet for several jobs
Building and running a crawler is expensive. Running 3 costs 3 times as much. An operator that needs pages for search, live answers and training therefore tends to fetch them once through shared infrastructure, then pass the result to the system that asked for it.
That is a sensible engineering choice, with no attempt to hide anything. But the decision about how to use the page happens inside the operator's systems, after the request reaches you. You cannot see that from your end.
What over 36% does to a rule written against a name
Most publishers who have acted on AI traffic have done so by company name. They found a company's user agent and either blocked it or allowed it.
With more than a third of automated activity coming from operators that use one fleet for several jobs, a company-level rule has wider effects than intended. Blocking the name may also block search indexing you wanted. Allowing it lets all the activities through. As far as we can see, changing how strict the rule is will not solve this. The rule needs to distinguish the activities.
What a mixed fleet might be doing on your site
A mixed fleet may carry out 4 activities on your site, each with a different value to you.
Search indexing helps readers find you. Grounding fetches a page for an assistant's live answer, sometimes with a citation and sometimes without. Agent retrieval carries out a person's request. Training collection copies content in bulk and returns nothing on the day it happens.
There is no reason to price all 4 the same way. Yet that is effectively what happens when you cannot distinguish them.
What it takes to tell them apart
You need to establish the operator using more than a self-declared label, classify the activity in commercially useful terms, and identify the requested URLs. The URLs show where interest concentrates. A total request count cannot do that.
Supertab Connect provides those 3 pieces of information from your own traffic at your existing CDN edge. They let you make a specific decision about an operator and an activity. A general decision about “AI” gives you very little to apply.
Where the figure comes from
The 36% figure comes from Cloudflare's July 2026 agentic internet bot report and covers its own network. It is the best public number we have. It will change, so check the date when you see it quoted, including here.
Your site's share will be different. Measure your own traffic before using a network-wide average to make decisions about your costs.