Skip to main content
A flood of machine request slips surrounding one verified human decision record

Your Dashboard Has a Bot Problem

When requests, crawlers, agents, and people share one chart, traffic can rise while human attention stays flat.

By Dellon S.13 min read

57.6%

Share of observed HTTP requests attributed to bots in the cited June 2026 Cloudflare data.

42.4%

Share attributed to humans in the same request-based view.

A marketing dashboard can be numerically correct and commercially wrong.

One human page load can create many requests. One crawler can fetch thousands of pages. A monitoring tool can make a useful request every minute. A malicious bot can mimic a browser. An AI agent can retrieve a product page for a real person and never send that person to the site.

When a report mixes those actors, traffic can rise while human attention stays flat. The number looks precise because the classification underneath it is missing.

Ask three questions before calling activity valuable: who acted, why it happened, and which commercial outcome followed.

The headline hides three audiences

“Bot traffic” sounds like one category. It contains search crawlers, AI training crawlers, real-time fetchers, shopping agents, monitoring services, security scanners, price scrapers, testing tools, and fraud systems.

In June 2026, Cloudflare Radar data summarized by IBM showed bots producing about 57.6% of observed HTTP requests and humans 42.4%. The request share is important for infrastructure and access policy. It cannot establish the composition of marketing sessions, exposed people, or customers by itself.

One crawler can retrieve a large archive in minutes. One person can create dozens of resource requests while loading a single page. An agent acting for one buyer can inspect many products. A ratio of requests therefore measures workload and access patterns before it measures audience.

Those actors should not share one marketing value. HUMAN Security's 2026 traffic report separates AI-driven automation into training crawlers, real-time scrapers, and agentic activity. Training crawlers accounted for about 67.5% of the AI automation it observed during 2025. Agentic activity was a much smaller share, reaching 1.7% by December, though it was growing quickly.

HUMAN also found that retail and ecommerce represented 62.5% of observed AI training crawler volume and 46.6% of agentic traffic. Those figures reflect its customer base. They still show why commerce teams need more detail than a single bot flag.

Human

Plausible attention or intent

Beneficial

Discovery, safety, or service

Operational

Monitoring, tests, and internal tools

Extractive

Content or price collection

Invalid

Manipulation, fraud, or abuse

A bot request can be valuable, costly, harmless, or fraudulent. Classification comes before valuation.

One dashboard, incompatible actors

Marketers combine ad-platform clicks and impressions, client-side analytics sessions, server logs, CDN requests, CRM records, ecommerce transactions, and sometimes a bot-management product. Each system sees a different event.

CDN

Requests before page load

Server

Routes and resources

Analytics

Client-side sessions

Ad platform

Filtered interactions

Revenue

Orders and contracts

Google Analytics automatically excludes known bots and spiders using Google research and the IAB International Spiders and Bots List. Google says users cannot turn that exclusion off or see how much known traffic was removed. That protects the report from obvious automation while creating an invisible boundary around the metric.

Google Ads applies separate invalid-click filtering. Its troubleshooting guide says Ads filters invalid clicks while Analytics can initially report sessions before potential filtering. It also explains that clicks, sessions, and users are different measures even when every event is human.

An invalid-click column confirms that filtering happened inside the ad platform. It does not tell the team how the same activity appeared in a CDN log, analytics property, CRM, or independent verification service. Each system applies different evidence at a different stage, so the same event can remain visible in one source and disappear from another.

  • Server requests can rise while sessions remain flat
  • Ad clicks can fall after invalid-traffic filtering
  • A crawler may never run the analytics tag
  • A native AI app may omit a conventional referrer
  • A person can block analytics and disappear
  • Repeated clicks can still form one session

Reconciliation fails when the team assumes every system is counting the same audience. The broader consequence appears in the AI search measurement gap.

Put the event definition beside every reported number. “Requests” should name the network layer and resource type. “Sessions” should name the analytics property and filtration. “Clicks” should state the platform and whether invalid interactions have already been removed. A shared label such as traffic hides more than it explains.

Mixed machine traffic records passing through a physical sorter into distinct evidence lanes
One report can contain infrastructure requests, platform clicks, client-side sessions, and verified outcomes. Sorting the unit is part of the analysis.

A bot visit can create value

Removing every automated event would produce a different kind of error. A search crawler can earn future discovery. An AI search fetcher can help a system cite the brand. A shopping agent can retrieve inventory for a person with real purchase intent. Monitoring can prevent revenue loss, and security automation can find risk.

The question is whether the automation created, protected, or extracted commercial value.

A useful exchange metric

Crawler requests ÷ referred requests

Cloudflare's crawl-to-referral ratio compares the amount of HTML content a platform retrieves with the visits it sends back.

In a June 2025 example, Cloudflare observed about 71,000 Claude crawler requests for each web referral it could identify. Cloudflare also explains the limit. Native apps may omit the Referer header, which can make the ratio look worse than the full reality. Crawling and referral behavior change quickly.

Referral counts also capture only one possible return. A system may cite a brand without a click, shape a later branded search, or help a user complete a task elsewhere. Those effects are harder to observe. Use referral data as a measured floor, then add citation sampling, declared-source surveys, controlled landing pages, or channel-specific transaction data where the business case justifies it.

Track discoverability, citations, qualified referrals, agent-assisted transactions, service outcomes, and infrastructure cost. Keep those outcomes in a machine-activity report instead of forcing them into a human-engagement chart.

The failure is classification

A pageview proves that an analytics system recorded a view event. It does not prove attention, memory, intent, or commercial value.

01

Actor

Human browser, verified crawler, declared AI bot, user-directed agent, internal system, monitoring service, unknown automation, or suspected malicious bot.

02

Purpose

Discovery, training, real-time retrieval, transaction, testing, monitoring, security, scraping, fraud, or unknown.

03

Outcome

No commercial signal, citation, referral, engaged session, qualified action, purchase, support resolution, cost, or blocked attempt.

Verified

Strong identity and behavioral evidence

Probable

Multiple signals support the classification

Unknown

Evidence is insufficient or contradictory

Store the rule, evidence, and timestamp behind each classification. Bot identities, operator ranges, browser signatures, and access patterns change. A metric that cannot reproduce its classification later cannot support a credible trend.

Do not force every event into human or bot. User-directed agents blur that line because software acts for a real person. Preserve both facts: the immediate actor was automated, while the beneficiary or decision-maker may have been human.

The unknown share is a metric. When it rises, the measurement surface is changing faster than the controls. This actor-purpose-outcome record gives agentic measurement a defensible unit of analysis.

Five physical evidence stages progressing from a translucent request to a fingerprinted order and parcel
Human confidence should increase as evidence moves from an infrastructure request toward an identified action and a commercial outcome.

Build a human-confidence layer

Start with the outcomes a person must create. A remembered name, product comparison, account action, qualified form, sales conversation, purchase, or repeat order is stronger evidence than a request alone.

Level 0

Request

The CDN or server received activity. No attention claim is allowed.

Level 1

Rendered interaction

Client code ran and recorded a plausible page or screen view.

Level 2

Meaningful behavior

Time, scroll, navigation, media use, or a similar pattern supports attention.

Level 3

Identified action

A person completes a verified account, form, trial, call, or qualified step.

Level 4

Commercial outcome

Revenue, margin, renewal, retention, or verified pipeline appears.

No level is perfect. Together they stop a Level 0 request from being presented as Level 4 demand.

Use multiple signals. User-agent labels can be spoofed. IP lists miss residential proxies. A single scroll event is easy to automate. Combine declared identity, verified operator ranges where available, request patterns, browser execution, interaction sequences, account state, server-side validation, and downstream outcomes. Apply privacy and consent rules to every signal collected.

Protect the actions that become business inputs. Forms should use server-side validation, rate limits, duplicate detection, and a qualified status from the receiving team. Ecommerce reports should account for cancellations, returns, and payment failure. Pipeline should distinguish a submitted lead from a sales-accepted opportunity.

Behavior is evidence, not identity. Long dwell time and deep scrolling can be automated. A short visit can still be valuable when a person finds the answer quickly. Use behavior to raise or lower confidence, then let verified downstream actions carry more weight.

Google's accredited measurement summary makes the limit explicit. Even platforms with continuous filtration cannot identify every invalid interaction proactively. The goal is a defensible estimate with disclosed methods.

Measure machine value separately

Once automated activity is separated, the team can evaluate it without pretending it is a person.

Access

Crawler requests, fetched pages, bytes served, origin load, blocks, and unknown automation.

Purpose

Training, search discovery, real-time retrieval, agentic action, monitoring, and suspicious behavior.

Return

Citations, referrals, assisted conversions, agent-completed transactions, and service outcomes.

Cost

Bandwidth, compute, paid APIs, content production, fraud review, and wasted lead handling.

A high crawler volume can be acceptable when it produces qualified discovery at a reasonable cost. A low-volume agent can be important if it assists transactions. A large crawl-to-referral gap can justify a new access policy. No universal threshold exists.

Make access decisions by declared purpose and observed behavior. A training crawler, answer-engine fetcher, and user-directed agent may come from the same operator under different identities. The site can choose different treatment for each, then watch whether the expected referral, citation, transaction, or service value appears.

Give finance the full exchange. Report the marginal cost of serving automated traffic, the share that reaches origin infrastructure, the content or inventory exposed, and the outcomes attributed with a stated confidence level. This keeps a high request count from being celebrated as demand or condemned as waste without evidence.

Akamai reported AI bot activity rising 300% in 2025 in publishing-focused research and warned about analytics pollution and infrastructure cost. That finding should trigger measurement work before it triggers a blanket block. The policy should reflect the value exchange and the site's business model.

Reconcile media to revenue

Marketing teams still need to buy media and explain performance. The bot problem makes reconciliation more important.

01

Platform delivery

Impressions and clicks

02

Reported invalid traffic

Filtered impressions and clicks

03

Analytics

Sessions and users

04

Human-confidence

Plausible attention and intent

05

Qualified action

Verified lead, trial, or purchase step

06

Commercial outcome

Revenue, pipeline, margin, or retention

Show both the count and the loss between each layer. Name the owner of every transition. A gap between ad clicks and sessions can come from filtering, consent, blocked JavaScript, page-load failure, redirects, or repeated clicks. A gap between sessions and qualified actions can come from targeting, friction, or automation that survived early filters.

Do not label every discrepancy as fraud. Diagnose it. For brand work, add direct traffic, branded search, survey-based awareness, qualified conversations, repeat visits, and later conversion among exposed groups. These measures have limits, but they are closer to the business question than raw request volume.

Build the bridge at a useful grain, such as campaign, landing route, country, device group, and week. Keep attribution windows and time zones consistent. Flag breaks in tagging or consent before interpreting the gap as audience quality. A reconciled loss table should explain where counts diverged and which team can investigate.

Use the bridge to change decisions. Media with strong platform delivery and weak verified outcomes may need a placement, audience, or fraud review. Content with heavy beneficial crawling and rising qualified referrals may deserve more investment. A dashboard becomes useful when it directs an owner toward a specific action.

A practical 30-day audit

The first month should create a trustworthy baseline, not a more elaborate dashboard.

Week 1

Map the systems

Record each source's event definition, filtration, time zone, attribution window, identity method, and owner.

Week 2

Classify actors

Sample CDN and server traffic, identify known categories, and preserve the unknown remainder.

Week 3

Connect outcomes

Build the confidence ladder through qualified action, refund, return, revenue, and margin.

Week 4

Change the report

Separate human-confidence, beneficial automation, unknown automation, and invalid traffic.

Document the sample size and the period for every actor estimate. Compare weekdays with weekends, launches with baseline weeks, and known campaign bursts with unexplained spikes. A single traffic snapshot can overstate a crawler surge or miss a recurring pattern.

Assign a review cadence. Fast-changing crawler identities may need weekly monitoring. Platform filtration and billing adjustments may need monthly reconciliation. The metric definition and control set should receive a formal quarterly review.

The new report may show less reach. That is progress when the remaining signal is easier to defend. It also gives micro-conversion measurement a cleaner baseline.

A clean human decision record separated from a closed stack of machine requests

Name the actor before the value.

The web can become more automated and still create more value for people. Marketing fails when it calls all of that activity human attention.

A request is infrastructure. A visit is an observation. Attention is an inference. Revenue is an outcome.

Ask which actor did what, for whom, and what changed afterward.

FAQs

Is most web traffic really bots?

Cloudflare data reported in June 2026 showed bots making roughly 57.6% of observed HTTP requests. That does not mean most people, sessions, pageviews, or buying decisions are bots. The result describes a specific request-based measurement.

Does Google Analytics remove bot traffic?

Google Analytics automatically excludes traffic from known bots and spiders using Google research and the IAB list. Google says users cannot disable the exclusion or see the excluded total. Unknown and sophisticated automation can still require additional analysis.

Should brands block all AI crawlers?

No universal answer fits every site. Separate training, search, real-time retrieval, agentic action, monitoring, and malicious traffic. Compare each category's discovery or service value with its infrastructure, content, and commercial cost.

How can marketers verify human traffic?

Use layered evidence including declared identity, operator validation, server patterns, browser execution, meaningful interaction, verified accounts, qualified actions, and downstream revenue. Preserve an unknown category instead of forcing weak signals into a human label.

What should replace raw traffic?

Keep raw traffic as an infrastructure metric. For marketing decisions, add human-confidence sessions, qualified actions, verified outcomes, machine-purpose measures, crawl-to-referral trends, and a reconciliation bridge from platform delivery to revenue.