Insights
September 14, 2026
to read

What is a mixed-use crawler?

A mixed-use crawler is a fleet that an operator uses for several jobs, including search indexing, live retrieval and training collection, through overlapping infrastructure. Cloudflare reported in July 2026 that mixed-use operators account for over 36% of automated activity. For that traffic, the company name alone does not explain the request's purpose. You may know who sent it and still misunderstand what you allowed.

Why one company runs one fleet for several jobs

Building and running a crawler is expensive. Running 3 costs 3 times as much. An operator that needs pages for search, live answers and training therefore tends to fetch them once through shared infrastructure, then pass the result to the system that asked for it.

That is a sensible engineering choice, with no attempt to hide anything. But the decision about how to use the page happens inside the operator's systems, after the request reaches you. You cannot see that from your end.

What over 36% does to a rule written against a name

Most publishers who have acted on AI traffic have done so by company name. They found a company's user agent and either blocked it or allowed it.

With more than a third of automated activity coming from operators that use one fleet for several jobs, a company-level rule has wider effects than intended. Blocking the name may also block search indexing you wanted. Allowing it lets all the activities through. As far as we can see, changing how strict the rule is will not solve this. The rule needs to distinguish the activities.

What a mixed fleet might be doing on your site

A mixed fleet may carry out 4 activities on your site, each with a different value to you.

Search indexing helps readers find you. Grounding fetches a page for an assistant's live answer, sometimes with a citation and sometimes without. Agent retrieval carries out a person's request. Training collection copies content in bulk and returns nothing on the day it happens.

There is no reason to price all 4 the same way. Yet that is effectively what happens when you cannot distinguish them.

What it takes to tell them apart

You need to establish the operator using more than a self-declared label, classify the activity in commercially useful terms, and identify the requested URLs. The URLs show where interest concentrates. A total request count cannot do that.

Supertab Connect provides those 3 pieces of information from your own traffic at your existing CDN edge. They let you make a specific decision about an operator and an activity. A general decision about “AI” gives you very little to apply.

Where the figure comes from

The 36% figure comes from Cloudflare's July 2026 agentic internet bot report and covers its own network. It is the best public number we have. It will change, so check the date when you see it quoted, including here.

Your site's share will be different. Measure your own traffic before using a network-wide average to make decisions about your costs.

Written by the Supertab Team

Pioneering the next generation of web monetization infrastructure and protocol-level content licensing.