Bots Block Bots: Why Anti-Bot Systems Flag LLM Browsers

02 October 2026 | 24 min read

LLM-driven browsers are browsers and fetchers that a language model operates on a person's behalf. Examples include Perplexity's Comet, ChatGPT's browsing, Claude computer use, and your own Playwright MCP agent. Anti-bot systems from Cloudflare, DataDome, Akamai, and HUMAN inspect and classify that traffic.

Bots block bots when legitimate agents look like scrapers, with datacenter IPs, headless fingerprints, machine-like input, and User-Agent strings that a site can't verify. Other blocks are the site owner's policy toward agents.

This guide measures those signals across common agent stacks and shows where legitimate fixes end and evasion begins.

Bots block bots: why anti-bot systems flag LLM-driven browsers

TL;DR

  • Production changes agent signals. Playwright MCP and Browser Use run headless by default on Linux servers without a display, so cloud deployments often combine HeadlessChrome with a datacenter IP.
  • Impersonating another agent failed. Akamai's and HUMAN's sites refused OpenAI's ChatGPT-User string from our machine 6 of 6 times.
  • Some blocks are policy. Since September 15, 2026, Cloudflare's preset for new ad-supported sites blocks agents on pages with ads, and verified agents aren't exempt.
  • Managed infrastructure passed the challenge. Every headless stack we tested failed a Cloudflare challenge, but ScrapingBee's Stealth proxy returned the real page 3 of 3 times.

What "bots block bots" means for LLM-driven browsers

An anti-bot system sits in front of a website and classifies each request as a person, allowed automation, or unwanted automation. It then serves, challenges, or blocks the request. LLM-driven browsers are one more kind of traffic to classify, and the browser-based ones run JavaScript and store cookies like any other browser.

Agent traffic comes from fetchers that an AI provider runs, browsers that a user runs, and agent stacks that you build. Most providers document how their fetchers identify themselves, and we captured the Claude-User string from a live request:

AgentWhat it sendsHow a site can verify itrobots.txt, per the vendor
ChatGPT-User (OpenAI)ChatGPT-User/1.0 tokenPublished IP listMay not apply to user-initiated visits
ChatGPT Work cloud browser (OpenAI's successor to Operator and ChatGPT agent)A Web Bot Auth signature for chatgpt.comSignature checkNot stated
Claude-User (Anthropic)Claude-User/1.0 tokenPublished IP listHonored
Perplexity-UserPerplexity-User/1.0 tokenPublished IP listGenerally ignored
Google-AgentGoogle-Agent tokenPublished IP list, plus signatures on some requestsGenerally ignored for user-triggered fetches
Comet, Claude in ChromeA Chrome identity with no documented agent tokenNo documented methodNot stated
Your Playwright MCP or Browser Use agentWhatever your Chromium build and host sendNothing, unless you signYour decision

An anti-bot system can't see intent directly. A scraper farm and a helpful agent can send identical signals:

Venn diagram where a scraper bot and an LLM agent share all five signal layers that a detector reads, so both get one risk score

Matching inputs produce a matching risk score.

Antoine Vastel of Castle argues the same in a bot-detection researcher's analysis of AI agents. Intent is hard to verify, so detection still depends on client, automation, infrastructure, and behavior signals.

The Signature Overlap Framework: five layers where agents look like scrapers

The Signature Overlap Framework groups what an anti-bot system reads into five layers, each with a fix that doesn't spoof an identity:

LayerWhat detectors checkWhat default agent stacks sentLegitimate fix
NetworkIP type, autonomous system number (ASN) reputation, geolocationAn IP address in a datacenter ASN, even from Claude's web fetch toolResidential or ISP exit IPs, with rate limits you'd accept on your own site
ProtocolTLS ClientHello (JA4), HTTP/2 settings, Client HintsLibrary TLS stacks, HTTP/1.1, a Chrome User-Agent string over non-Chrome TLSA full browser engine throughout the request path
Environmentnavigator.webdriver, Chrome DevTools Protocol (CDP) traces, headless brands, WebGLwebdriver set to true, HeadlessChrome, a SwiftShader software rendererHeaded, full browsers, plus the network fix. Do-it-yourself patches can add inconsistencies
BehaviorPointer paths, key timing, isTrusted, pacingOne direct pointer movement per click, keystrokes about 8 ms apart, synthetic eventsFewer UI actions, caching, and pacing, never faked mouse movement
IdentityUser-Agent tokens, IP lists, reverse DNS, signaturesA default Chrome User-Agent string or an unsigned tokenYour own token plus Web Bot Auth, never another company's token

Our measurements come from one Mac on a residential connection and one cloud server. Each result is an observation of a named configuration, not a rate, so you can rerun it on your own stack.

Network layer

A laptop test on a home connection doesn't reproduce the network of a typical cloud deployment, because its traffic comes from a residential IP. An agent deployed in a cloud region sends a datacenter IP unless it routes through a proxy. Even Claude's web fetch tool reached us from AS396982, Google Cloud's network.

We ran our agents on a GitHub Actions server in Microsoft's Azure network. Every cloud-server run stopped at a Cloudflare challenge. From a residential connection, headed Browser Use passed the same challenge.

A University of Bamberg crawl of the Tranco top 10,000 sites ran from a Hetzner datacenter. It measured how often default Playwright browsers received status codes associated with blocking. Headless Chromium received a 403, 429, or 503 on 15.2% of sites, and headed Chromium and Firefox received one on 6.8% to 7.2%.

Protocol layer

The TLS handshake fingerprints the HTTP client before any page loads. Its JA4 fingerprint, a short code built from the handshake, was t13d1812h1 for Python's requests and t13d1517h2 for Chromium 153. Chromium 153 also negotiated the post-quantum key exchange X25519MLKEM768, which Cloudflare's /cdn-cgi/trace shows in an undocumented kex field, while requests and curl used plain X25519.

Playwright's headless shell changes the handshake as well. The chromium-headless-shell build that Playwright uses for headless=True sent 16 TLS extensions, and full Chromium sent 17. It was missing 0xca34, BoringSSL's trust anchors extension.

A detector can compare each value with the browser that the User-Agent string claims. A Chrome User-Agent string sent over a Python handshake fails that comparison.

Environment layer

Playwright 1.63 keeps navigator.webdriver set to true in all three modes we tested, and its own test suite asserts that value. Playwright MCP 0.0.82 and Browser Use 0.13.10 add --disable-blink-features=AutomationControlled by default to hide that flag.

A comment in the Browser Use launch profile states the goal as "we mask the automation fingerprint via JS and other flags." When you don't set a headless option, both frameworks switch to headless on a Linux server without a display, which is a common production setup. The Playwright MCP config sets this default in code, and Browser Use does the same.

The fingerprint test on deviceandbrowserinfo.com produced different results for two default stacks:

Fingerprint test results, You are a bot for headless Playwright 1.63 and You are human for headed Browser Use 0.13.10

Same page, same machine, one visit each.

The test also reports which checks triggered for each stack:

Stack and modeClassified asChecks that triggered
Playwright 1.63, headless shellBotHeadlessChrome User-Agent, webdriver, CDP control, timing
Playwright 1.63, headedBotwebdriver, CDP control, timing
Playwright MCP 0.0.82, headlessBotHeadlessChrome User-Agent, CDP control, timing
Playwright MCP 0.0.82, headed (default)BotCDP control
Browser Use 0.13.10, headlessBotHeadlessChrome User-Agent
Browser Use 0.13.10, headed (default)HumanNone

The masking frameworks hide webdriver but keep HeadlessChrome in the User-Agent string, so a headless run sends inconsistent signals. The headless shell also sends HeadlessChrome in its sec-ch-ua header and renders WebGL through SwiftShader. An Inria study found that stealth and anti-detection add-ons often make agents easier to detect. For background on these modes, read what a headless browser is.

The Bamberg crawl found that 34% of sites read navigator.webdriver. Of 784 re-crawled sites that sent a 403 only to headless Chromium, 590 served the page without HeadlessChrome in the User-Agent string and Client Hints.

Behavior layer

We asked Browser Use, the framework from our AI browser automation tutorial, to type a phrase into a search box and press a button. The page logged every input event:

Timeline of one agent task with 4.7 s of model thinking and no input, then 15 keystrokes a median 8 ms apart, one direct pointer movement, and a click

The model's thinking time looks human, but the input events after it look machine-like.

We ran the task headless three times with Claude Sonnet 5 and twice with GPT-5 mini. Every run produced one mousemove and a click at the same sub-pixel offset from the button's center. Median inter-keystroke intervals were 8 to 9 ms, and each run fired the same three untrusted events, two input and one blur. Only the idle time before the first action varied, from 4.7 to 6.8 s.

The model picks the action, and the framework produces the input events. We recorded these traces on our own test page, so they show exactly what the agent sends, without any site's defenses affecting them.

We set up Playwright MCP as in our tutorial on using Playwright with an MCP server, and it left a different but also machine-like trace. Its default browser_type call fired one insertText input event that contained the whole phrase, and no keystroke events. With slowly: true, it typed at a median of 0.5 ms per key.

Detection research has measured this pattern, and DataDome's own block page lists clicking faster than a human as one possible cause. Akamai's researchers found that more than 98% of requests from agentic browsers such as Comet and Atlas carried little or no mouse movement. Akamai's legacy mouse models flagged under 1% of those requests, because there was almost nothing to evaluate. The FP-Agent study used behavior alone to distinguish seven agents from people, with an F1 score of 0.9994.

Cloudflare launched Precursor in July 2026 for Enterprise Bot Management. It collects behavioral signals across a whole session to separate human from agentic traffic.

Identity layer

Identity is the one layer that you declare, and a declaration matters most when a site can check it. Cloudflare's managed rules flag a request as a fake bot when its User-Agent string matches a known bot but its source can't be verified. Akamai's Bot Manager has a detection category named Impersonators of Known Bots. Our guide on how websites use User-Agent strings to detect bots covers the basics.

Why Cloudflare blocks AI agents: the Search, Agent, and Training controls

In July 2026, Cloudflare split AI traffic into three behaviors. Search collects content to answer questions later, and Training uses content to train a model. Agent acts in real time for a person, and Cloudflare's examples include ChatGPT-User and browser-using agents such as Gemini or Claude driving Chrome.

On September 15, 2026, Cloudflare changed the defaults for new domains. Onboarding offers one of two presets, depending on whether the site earns money from ads:

SettingSite without adsSite with ads
SearchAllowAllow
TrainingAllowDisallow AI Training
AgentAllowBlock on pages with ads

Cloudflare also changed settings for existing domains. If a domain had the older Block AI Bots toggle enabled, Cloudflare migrated it to block the Agent category on pages with ads. Cloudflare's reasoning is that an agent fetches a page with nobody there to see the ads. Owners can change the controls on any plan at any time.

The older controls were already widely used. Cloudflare told WIRED that it blocked 416 billion AI bot requests between July 1 and early December 2025.

On Cloudflare, verification no longer means permission. Since July 2026, Cloudflare allows verified-bot traffic only when the site's settings permit its category. So a verified agent is still blocked on a page where the Agent category is blocked. The block comes from the site's settings, so ask the owner to allow your agent instead of bypassing the block.

First-hand test: what protected pages return to AI agents

We sent default agent stacks to a page behind a Cloudflare Managed Challenge, and we varied the User-Agent string. We also fetched the page through ScrapingBee's Stealth proxy, a proxy pool built for the hardest sites to scrape.

Our stacks were Playwright 1.63 and Browser Use 0.13.10, both on Chromium 153, on one residential connection and one Azure cloud server. We made one attempt with each browser configuration per run and waited 15 s for the managed challenge to complete by itself. ScrapingBee ran in its default mode, which retries a failed request on other proxies for up to 140 s. To keep the comparison fair, every stack that failed got a rerun with the same 140-second timeout, and no run clicked or solved a challenge.

What a Cloudflare Managed Challenge returned to each stack

The target was a public practice page whose operator runs a Cloudflare Managed Challenge for testing scrapers. Every Playwright and Python client got the challenge interstitial as its first response:

Cloudflare interstitial reading Performing security verification, with a Verify you are human checkbox, captured from a headed Playwright browser

Headed Playwright got an interactive checkbox, which we left unclicked. The site's domain is cropped out.

The first response contained Cloudflare's challenge headers. Its Critical-CH header asks for 17 client hints, including the browser's full version list, architecture, and platform version, shortened here:

HTTP/2 403
server: cloudflare
cf-mitigated: challenge
critical-ch: Sec-CH-UA-Bitness, Sec-CH-UA-Arch, Sec-CH-UA-Full-Version,
             Sec-CH-UA-Mobile, Sec-CH-UA-Model, Sec-CH-UA-Platform-Version,
             Sec-CH-UA-Full-Version-List
server-timing: chlray;desc="a40a6f86fb34a55f"
cf-ray: a40a6f86fb34a55f-MRS

The page's window._cf_chl_opt object sets cType: 'managed', which confirms the challenge type. The number of attempts varied by configuration, from one to twelve, so read each row as a count, not a rate:

ClientPage servedWhat came back
Python requests, default or Chrome User-Agent string0 of 3 each403 with cf-mitigated: challenge
Playwright headless shell or new headless0 of 3 eachThe interstitial stayed
Playwright headed, with any of 3 User-Agent strings0 of 12The interstitial stayed
Browser Use, headless0 of 3The interstitial stayed
Browser Use, headed, home connection5 of 6The real page within 15 s, first response not recorded
Playwright or Browser Use, headed, on an Azure cloud server0 of 3 eachThe interstitial stayed for 140 s
Claude's web fetch tool (Claude-User, Google Cloud IP)0 of 1Tool error url_not_accessible, inconclusive
ScrapingBee Stealth proxy3 of 3The real page

In this small test, changing the User-Agent string did not change the result, but the outcome varied by browser setup. On the home connection, headed Browser Use reached the real page 5 of 6 times, including runs where it went first. Headed Playwright never passed the interstitial, even in 140 s.

From the cloud server, headed Browser Use failed too, so a result from home may not hold in production. Each setup differs in several signals at once, so these runs show which setups passed, not which signal decided the result.

We also tested a second Cloudflare-fronted page that did not apply a challenge to our requests. It served Python requests, curl, and headless Playwright with HTTP 200. The result depends on each site's rules and each client's signals, not on one Cloudflare rule against agents.

Impersonated, declared, and default identities

The User-Agent string didn't change the managed challenge result, so we tested identity against the anti-bot vendors' own public homepages. DataDome's homepage refused each local stack in its single test attempt, including headed Browser Use, so the table covers Akamai and HUMAN. In September 2026, the same headed Playwright browser sent each of three User-Agent strings three times, rotating their order:

User-Agent stringAkamai homepageHUMAN homepage
Default Chromium string3 of 3 served2 of 3 served
Default string plus a declared agent token0 of 33 of 3
OpenAI's ChatGPT-User string0 of 30 of 3

The impersonated identity failed every time, even from a browser that both sites served with its default string. Requests using OpenAI's string gained no access, whether a fake-bot rule or a policy against AI agents refused them.

Sent from outside OpenAI's published IP ranges, that string is exactly the claim that verification is designed to catch. We used it only to measure that outcome.

The declared token got different results. Akamai's site refused it 3 of 3 times, and HUMAN's site accepted it 3 of 3 times. Treat Akamai's refusal as a policy to respect. HUMAN's homepage also runs behind Cloudflare's CDN, and a vendor's own site shows one configuration, so test the sites your agent actually visits.

How to identify the block

The response can show the type of block and sometimes suggest which layer contributed to it. Check it before you change anything:

What you seeVendorWhat it meansWhat to do
403, cf-mitigated: challenge, cType: 'managed'CloudflareA rule or bot score triggered a browser checkConfirm robots.txt and terms permit access, then use supported browser and network infrastructure, or stop
Error 1020, or a "you have been blocked" pageCloudflareA firewall rule matchedRespect it, or ask the owner
402 with a crawler-price headerCloudflarePay per crawl sets a price for the pagePay, or skip the page
403 "You have been blocked", with an x-datadome headerDataDomeBot protection refused the requestConfirm robots.txt and terms permit access, then use supported browser and network infrastructure, or stop
403 "Access Denied" with a reference numberAkamaiBot Manager refused the requestConfirm robots.txt and terms permit access, then use supported browser and network infrastructure, or stop
403 "Access to this page has been denied", with a press-and-hold checkHUMAN (PerimeterX)Bot protection refused the requestConfirm robots.txt and terms permit access, then use supported browser and network infrastructure, or stop

In a browser, HUMAN and DataDome also show a page that's easy to recognize in your agent's screenshots:

Block pages served to headless Playwright, a press-and-hold check from HUMAN and a You have been blocked page from DataDome

Default headless Playwright got these pages on HUMAN's and DataDome's own homepages in September 2026.

Cloudflare documents how to detect a challenge response through the cf-mitigated header, which is a check to add to your fetch code.

When the response doesn't name the layer, change one thing at a time. Rerun the same request through a residential exit IP, then in a headed browser. If one change alters the result, that layer likely contributed to the block. Do this for a heuristic challenge, never for a policy block.

The legitimate-use line: what counts as a fair fix and what counts as evasion

The line separates signals that a site owner sets from signals that a detector infers. Owners set policy with robots.txt, Content Signals, TDMRep and AI.txt, Cloudflare's Agent and Training controls, pay-per-crawl prices, login walls, and terms of use. Detectors infer automation from datacenter ranges, headless traces, protocol mismatches, and machine timing.

Some policies apply only to agents that a detector can classify. In the FP-Agent study, Cloudflare's free AI bot controls fully blocked only 1 of 7 agents, the one that identifies itself as a verified bot. An agent that looks like an ordinary browser may never receive the owner's decision. Masking your agent's signals can therefore hide it from the policies a detector enforces, which is a strong reason to declare who you are.

These rules hold at every layer:

  • Never automate CAPTCHA solving. A CAPTCHA is a check for a human, and Anthropic says its bots don't bypass CAPTCHAs.
  • Never impersonate another company's bot. Googlebot and ChatGPT-User strings belong to their owners, and fake-bot rules flag them when the source can't be verified.
  • Never collect content behind a login. ScrapingBee's acceptable use policy forbids collecting non-public information and names data behind a login as an example.
  • Check visible policy before each fetch. Read robots.txt for your own token and for the default * rules, and skip pages whose Content-Signal line sets ai-input=no.
  • Treat an explicit refusal as final. Do not retry a policy 403, an Agent block, or a robots.txt Disallow under another identity or proxy.

Robots.txt is less clear for agents than for crawlers. OpenAI's docs say of ChatGPT-User, "Because these actions are initiated by a user, robots.txt rules may not apply." Perplexity and Google say their user-triggered fetchers generally ignore it, while Anthropic says Claude-User honors it. Cloudflare's September post adds that no established directive exists yet for stating preferences to agents.

The closest signal is ai-input in Content Signals, which covers content given to AI models for real-time answers. Cloudflare's own robots.txt sets it like this:

Content-Signal: ai-train=yes, search=yes, ai-input=yes

The reference MCP fetch server is cautious. It checks robots.txt before autonomous fetches and sends a declared ModelContextProtocol/1.0 token. Our guide to robots.txt for scrapers covers the parsing.

Web Bot Auth: identity that a site can verify

Web Bot Auth replaces a claimed identity with a signed one. The agent signs each request with an Ed25519 key under HTTP Message Signatures, defined in RFC 9421. A Signature-Agent header points to a public key directory at /.well-known/http-message-signatures-directory. ChatGPT Work's cloud browser signs this way, and Google signs part of its Google-Agent traffic.

For Cloudflare, you generate a key, host the signed directory, and register the bot through the dashboard's Bot Submission Form. This example requires the requests and cryptography packages. It uses the RFC 9421 test key and gets a valid result from Cloudflare's research verifier:

import base64
import hashlib
import json
import os
import time
from urllib.parse import urlparse

import requests
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey

# RFC 9421 test key, as published in cloudflare/web-bot-auth. Never use it in production.
JWK = {"kty": "OKP", "crv": "Ed25519",
       "d": "n4Ni-HpISpVObnQMW0wOhCKROaIKqKtW_2ZYb2p9KcU",
       "x": "JrQLj5P_89iXES9-vFgrIy29clF9CC_oPPsw3c5D0bs"}
AGENT = "https://http-message-signatures-example.research.cloudflare.com"


def b64url(data: bytes) -> str:
    return base64.urlsafe_b64encode(data).rstrip(b"=").decode()


def sign(url: str) -> dict:
    key = Ed25519PrivateKey.from_private_bytes(
        base64.urlsafe_b64decode(JWK["d"] + "="))
    # keyid is the RFC 7638 thumbprint of the public key
    canon = json.dumps({k: JWK[k] for k in ("crv", "kty", "x")},
                       separators=(",", ":"))
    keyid = b64url(hashlib.sha256(canon.encode()).digest())
    now = int(time.time())
    params = (f'("@authority" "signature-agent");created={now};'
              f'expires={now + 60};keyid="{keyid}";alg="ed25519";'
              f'nonce="{b64url(os.urandom(32))}";tag="web-bot-auth"')
    # Signature base: one line per covered component, then the parameters
    base = (f'"@authority": {urlparse(url).netloc}\n'
            f'"signature-agent": "{AGENT}"\n'
            f'"@signature-params": {params}')
    signature = base64.b64encode(key.sign(base.encode())).decode()
    return {"Signature-Agent": f'"{AGENT}"',
            "Signature-Input": f"sig1={params}",
            "Signature": f"sig1=:{signature}:"}


url = f"{AGENT}/v0/api/verify"
print("unsigned:", requests.get(url, timeout=30).text)
print("signed:  ", requests.get(url, headers=sign(url), timeout=30).text)

The verifier returns neutral without a signature and valid with one:

unsigned: neutral
signed:   valid

The IETF Internet-Draft draft-ietf-webbotauth-httpsig-protocol-00 specifies Signature-Agent as a dictionary. Cloudflare's July 2026 docs and its verifier reject that form, so send Cloudflare the quoted-string form.

Cloudflare's docs also say it is experimenting with the Forwarded header from RFC 7239. With it, an operator's identity can stay attached when requests pass through intermediaries that Cloudflare trusts.

On Cloudflare, a signature lets a site that checks signatures apply its policy to you, not to every unverified bot. It doesn't remove a category block.

Managed infrastructure that reduces false-positive blocks

ScrapingBee's HTML API changes the network, protocol, and environment layers for you. It runs JavaScript rendering in its own Chromium browsers and sends requests through rotating proxy pools. It returns a rendered page per request, which works well for agents that read and extract.

To try the Stealth proxy before writing code, select it in the dashboard's HTML Playground:

ScrapingBee HTML Playground with example.com as the URL, Auto mode off, Stealth Proxy selected, and JavaScript Rendering on

You pay 75 credits only when a Stealth request succeeds.

In 12 JavaScript-rendered samples, every request had a Chromium TLS fingerprint and Chrome's exact HTTP/2 fingerprint under a Chromium-family User-Agent string. All three Premium samples came from ISP networks, including Comcast. This result matches the network-layer fix in the Signature Overlap Framework.

Auto-Mode picks the cheapest configuration that works and bills only that one. Use it with a policy check, so your agent skips pages where the site's stated policy says no.

ScrapingBee fetches the page with its own browser identity. The preflight below is conservative. It requires the rules for your agent's token and the default * rules to permit access. It can't verify which rules apply to the browser identity that fetches the page.

The fetch limits the escalation to the Stealth proxy and requires Python 3.10+, the requests package, and your API key in SCRAPINGBEE_API_KEY:

import os
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser

import requests

API = "https://app.scrapingbee.com/api/v1/"
AGENT = "ExampleResearchAgent"  # your own product token


def policy_allows(url: str) -> bool:
    origin = "{0.scheme}://{0.netloc}".format(urlparse(url))
    try:
        robots = requests.get(f"{origin}/robots.txt", timeout=15,
                              headers={"User-Agent": f"{AGENT}/1.0"})
    except requests.RequestException:
        return False  # can't read the policy, so don't fetch
    if robots.status_code in (404, 410):
        return True  # no robots.txt file is published
    if robots.status_code >= 400:
        return False  # 401, 403, 429, 451, 5xx: policy unknown, so don't fetch
    lines = robots.text.splitlines()
    parser = RobotFileParser()
    parser.parse(lines)
    opted_out = any(line.lower().replace(" ", "").startswith("content-signal:")
                    and "ai-input=no" in line.lower().replace(" ", "")
                    for line in lines)
    # Conservative: also require the default rules, since ScrapingBee uses its own browser identity
    return (parser.can_fetch(AGENT, url) and parser.can_fetch("*", url)
            and not opted_out)


def fetch(url: str) -> str | None:
    if not policy_allows(url):
        print(url, "skipped: the site's stated policy says no")
        return None
    response = requests.get(
        API,
        headers={"Authorization": f"Bearer {os.environ['SCRAPINGBEE_API_KEY']}"},
        params={"url": url, "mode": "auto", "max_cost": "75"},
        timeout=180,
    )
    response.raise_for_status()  # a 500 means every configuration failed, unbilled
    print(url, "Spb-auto-cost:", response.headers.get("Spb-auto-cost"))
    return response.text


for page in ("https://www.scrapingbee.com/blog/",
             "https://www.google.com/search?q=web+bot+auth"):
    fetch(page)

The blog page stopped at the cheapest configuration, and robots.txt disallowed the search page:

https://www.scrapingbee.com/blog/ Spb-auto-cost: 1
https://www.google.com/search?q=web+bot+auth skipped: the site's stated policy says no

On the Cloudflare practice page, Auto-Mode escalated to the Stealth proxy and returned the real page at 75 credits. If your agent's fetch tool accepts an HTTP proxy, Proxy Mode gives the same access through a standard proxy setting, without HTML API calls.

For agents that support MCP, ScrapingBee's MCP server exposes the same fetches as tools. In September 2026, the live endpoint listed 18 tools, from page fetches to Amazon and YouTube data. It accepts the API key as a Bearer header, which keeps the key out of the URL, as in this Cursor mcp.json entry:

{
  "mcpServers": {
    "scrapingbee": {
      "url": "https://mcp.scrapingbee.com/mcp",
      "headers": { "Authorization": "Bearer YOUR_SCRAPINGBEE_API_KEY" }
    }
  }
}

Each tool call costs the same as the endpoint it wraps. The page tools use Auto-Mode by default, and max_cost sets a credit limit per call. Run your policy check before the agent calls them.

Extraction can reduce the returned payload. In our comparison of MCP servers for web scraping, ScrapingBee's extract_rules returned 623 tokens for one catalog page, compared with 9,673 tokens of raw HTML. Both calls cost 1 credit.

Reading fetched pages also reduces behavior signals, because your agent performs fewer UI actions for a detector to measure. The identity layer is your agent's responsibility. If a site accepts only verified agents, use Web Bot Auth through a request path that preserves the signed headers and components.

Cache fetched pages where freshness allows, so your agent avoids repeat requests.

What it costs: credits per request and a cost example

ScrapingBee bills credits per successful request, and the proxy option sets the rate:

ProxyWithout JavaScriptWith JavaScript
Classic15
Premium1025
StealthRuns with JavaScript75

AI extraction adds 5 credits to each request that uses it. Even if every page needs the Stealth proxy with extraction, 10,000 pages a month use 800,000 credits. The 800,000 total is within the Startup plan's 1 million monthly credits, which ScrapingBee's pricing page listed at $99 in September 2026. Per agent task, 20 such pages cost 1,600 credits.

Auto-Mode lowers those totals when some pages succeed on Classic at 1 or 5 credits. Failed requests that return 500 cost nothing.

Final thoughts

"Bots block bots" starts as a signal problem, because the agent stacks we tested looked like scrapers on several of the five layers. Test from the network you deploy on, because a fix that passed at home failed from our cloud server. You can handle a heuristic challenge with an infrastructure fix, but a policy block such as Cloudflare's Agent setting is the owner's decision.

Masking your agent can also hide it from that policy, so check robots.txt and Content Signals first, and sign requests where sites accept signatures. ScrapingBee's free trial includes 1,000 credits, enough to test your blocked URL with a policy-checked Auto-Mode fetch.

Bots block bots FAQs

Are AI agents considered bots?

Yes. Anti-bot systems usually treat AI agents as automated traffic, then apply the site's rules to them. Declared agents such as ChatGPT-User and Claude-User can be checked against published IP lists or signatures. An agent that impersonates another company's bot fails that check.

Does Cloudflare block AI bots?

Yes, when the site's Cloudflare settings block that traffic category. Cloudflare's controls treat Search, Agent, and Training traffic separately. Since September 15, 2026, its preset for new ad-supported domains blocks agents on pages with ads. Verified agents are blocked too when the site blocks their category.

GPTBot vs. ChatGPT-User: what's the difference?

GPTBot is OpenAI's crawler for training data, and ChatGPT-User makes visits that a person starts in ChatGPT. The two have separate tokens. OpenAI documents robots.txt controls for GPTBot. For ChatGPT-User, it says the agent doesn't crawl automatically and that robots.txt rules may not apply.

Does robots.txt stop AI crawlers and agents?

Only the ones that honor it, because robots.txt is a published request, not access control. Crawlers such as GPTBot say they follow it. For user-triggered agents, Anthropic says Claude-User honors it, OpenAI says it may not apply to ChatGPT-User, and Perplexity and Google say theirs generally ignore it.

Can AI agents solve CAPTCHAs?

A legitimate agent shouldn't try. A CAPTCHA exists to confirm that a human is present, so automated solving is an evasion technique. When an agent meets a CAPTCHA, it should stop, ask its user to take over where appropriate, or skip the page. Anthropic, for example, says its bots don't try to bypass CAPTCHAs.

It depends on jurisdiction, the site's terms, and the data. This isn't legal advice. Public pages carry less risk than logged-in accounts, which were key in Amazon's case against Perplexity. In August 2026, the Ninth Circuit vacated the preliminary injunction and remanded the case without a final ruling on legality.

How much does ScrapingBee cost?

ScrapingBee bills credits per successful request, from 1 credit for a Classic request without JavaScript to 75 for the Stealth proxy. Failed requests that return 500 cost nothing. In September 2026, the Startup plan listed 1 million monthly credits at $99, and the free trial includes 1,000 credits.

image description
Satyam Tripathi

Satyam works in developer marketing for companies building web data and AI infrastructure products. He creates technical and product content that helps developers understand and use complex technologies.

Auto-mode picks the configuration that successfully scrapes your page

Try it now