A small philanthropy database in Austin, Texas, has revealed an in depth account of a 12 months spent preventing automated site visitors, reporting that bots accounted for greater than 99% of requests hitting its servers and that one AI search crawler learn its pages 35,000 instances for each human customer it despatched again.
Nick Grey, founding father of PatronView, a analysis database protecting American museum and cultural establishment donors, revealed the findings on August 7, 2026, in a put up titled “99% of My Web site Visitors Is Bots.” The piece walks by a 12 months of server logs, firewall adjustments, and crawler conduct throughout a 1.5 million web page website constructed from IRS 990 varieties, donor partitions, and annual reviews.
Grey posted a abstract on X the identical day, writing that he was “shocked to see that 99% of my site visitors is from bots” and that Anthropic‘s crawler “scrapes my website 35,000 instances for each 1 customer they ship.” The put up drew replies from different website operators describing related experiences, together with Arvid Kahl of Podscan, who mentioned most of his operational work now includes stopping his servers from being overwhelmed by bot site visitors.
A server log tells a unique story than an analytics dashboard
The core measurement separates what customer analytics instruments report from what a server really solutions. Within the week Grey revealed the piece, PatronView’s server answered 2.5 million requests and served 1.28 million full pages, whereas the location’s self-hosted Believable analytics recorded 5,977 human pageviews, a ratio of roughly 214 non-human web page masses for each web page load counted as human.
Grey attributes the hole to how JavaScript-based analytics instruments work. Believable, like Fathom or Google Analytics, solely counts guests whose browsers execute JavaScript, and virtually no bot does, so a website proprietor checking solely that sort of dashboard has no visibility into what is definitely putting the server.
The location itself is constructed by scraping public philanthropic disclosure paperwork, a element Grey addresses instantly, distinguishing his personal periodic assortment of supply paperwork from the 1000’s of automated requests his personal website receives each day.
Nationwide site visitors spikes preceded the crawler findings
Earlier than Grey remoted particular person AI crawlers, the location skilled site visitors occasions tied to geography fairly than a single firm. In November 2025, roughly 4 thousand periods appeared over a couple of days, every visiting precisely one web page with a 99% bounce price and no referrer, targeting fund pages that ordinarily draw solely about 10% of actual customer site visitors.
The bigger spike got here on April 22, 2026, when the location recorded 3.6 million requests in a single day, originating from 361,844 distinctive IP addresses, the massive majority positioned in China. Cloudflare‘s Managed Problem system, described within the put up as an invisible CAPTCHA, absorbed 1.18 million of these requests inside the first ten hours, although a notable share was passing the problem fairly than failing it.
Grey responded by blocking the whole nation of China on the community edge on April 23, then prolonged the identical block to Vietnam and Singapore after related patterns appeared from these international locations. He helps the choice together with his personal viewers knowledge: PatronView’s real search site visitors runs 95.9% United States and 1.3% Canada, for a database protecting American donors revealed in English.
Replies to Grey’s public put up concerning the incident, cited in his write-up, counsel the sample is widespread amongst unbiased website operators. Matt Paulson of MarketBeat really useful including Russia to any block checklist. Jack Ellis of Fathom Analytics reported that clients had seen a comparable wave of Chinese language-origin spam site visitors over roughly six months earlier than it dropped off fully. Jeremy Brandt described his personal nation block checklist as “intensive.” Rodrigo Rocco flagged a more durable drawback that recurs later in Grey’s findings: site visitors more and more arrives by 1000’s of residential IP addresses making single calls every, defeating geography-based and volume-based blocking concurrently.
Measuring the crawl-to-referral ratio instantly
Essentially the most particular figures in Grey’s put up concern AI crawlers constructed to energy search and reply merchandise fairly than general-purpose coaching crawlers. Cloudflare has beforehand acknowledged, in figures Grey cites, that Anthropic’s crawlers run at roughly 3,000 pages crawled for each one customer referred. Grey’s personal measurement, taken in June 2026 utilizing Cloudflare’s AI crawler dashboard, put the determine for his website at 35,000 to 1.
The underlying numbers got here from two separate consumer brokers Anthropic operates. Claude-SearchBot, the crawler Anthropic makes use of to construct search outcomes, requested 420,680 pages from PatronView over one week. In that very same week, Claude-Person, the distinct consumer agent Anthropic sends when somebody utilizing Claude asks it to fetch a selected web page on an individual’s behalf, delivered 12 human guests to the location. Bandwidth informed the identical story from a unique angle: the location served 4.63 gigabytes to the search crawler that week and 175 kilobytes to the people it referred.
Grey disclosed that PatronView was itself constructed with Claude Code and that he continues to make use of Claude Code to refine his Cloudflare safety guidelines, a element he presents as an acknowledged irony, earlier than describing his choice to dam Claude-SearchBot on the firewall stage. Following the block, the crawler’s request quantity fell from roughly 60,000 requests a day to about 25 makes an attempt a day, a sample he learn as proof that Anthropic’s crawler respects a typical HTTP 403 access-denied response fairly than trying to bypass it.
That measurement produced a broader metric Grey now applies to each crawler hitting the location: pages crawled per human customer referred. Googlebot measured at 46 crawls per customer, a ratio Grey describes as incomes its hold. Bingbot measured at 406 to 1, worse however nonetheless defensible in his framing. The AI crawlers constructed for coaching giant language fashions measured far larger, with Claude-SearchBot at 35,000 to 1 earlier than the block.
A second crawler with no referral path in any respect
Whereas getting ready the put up, Grey recognized Amzn-SearchBot, the crawler Amazon operates to feed its Rufus procuring assistant and Alexa’s reply merchandise, as the location’s highest-volume crawler at roughly 117,000 requests per day, practically double what Claude-SearchBot had reached in June earlier than its personal block took impact. As a result of the crawler exists to generate solutions inside Amazon’s personal merchandise fairly than to ship site visitors elsewhere, Grey wrote that it might by no means refer guests again and that he doubted it might even present attribution. He blocked it two days earlier than publishing the put up, describing the change as taking two minutes to implement.
Grey additionally credit Cloudflare’s managed AI Crawl Management function with blocking declared coaching crawlers, together with GPTBot, ClaudeBot, CCBot, and Bytespider, earlier than his personal customized guidelines run. The identical dashboard confirmed Bingbot requesting 158,610 pages in opposition to 680 referred guests, a ratio Grey characterised as acceptable provided that his Bing referrals proceed to develop.
A CAPTCHA function that value greater than the location it protected
Individually from the crawler-blocking work, Grey described discovering that Cloudflare’s JavaScript Detections function, which had been injecting a problem script into each web page load all year long, was consuming 2,875 milliseconds of load time on a mid-range cellphone in opposition to 278 milliseconds for the location’s personal JavaScript. He recognized the script as the first motive his cellular Lighthouse efficiency rating sat at 58, noting that the function’s diagnostic output was not readable by his firewall rule interface.
Grey disabled the function on August 5, 2026, and reported a Lighthouse rating of 99 inside an hour. 5 hours after the change, a scraper working from Microsoft Azure retrieved 23,000 pages in a single hour, spreading requests throughout greater than 80 IP addresses that every stayed below Grey’s present price restrict, proof he learn as exhibiting the disabled function had actually been suppressing that site visitors. A separate wave of residential IP addresses started presenting as Chrome variations 118 by 120, browser releases from 2023, following the identical disabling.
Residential IP site visitors defeats geography and volume-based guidelines
Grey related the August sample to a structurally related occasion from November 2025, wherein site visitors arrived from roughly forty international locations in near-identical numbers on the identical day, a distribution he characterizes because the signature of a rented residential proxy community wherein 1000’s of unusual residence web connections are used, by the request, to distribute a single buyer’s site visitors so it seems to originate from actual particular person customers.
By late July 2026, the sample reappeared, targeting American residential web connections fairly than distributed internationally. Distinctive IP addresses hitting the location climbed from a baseline of roughly 18,000 per day to 124,000 on July 31, 2026, whereas nothing concerning the underlying website had modified throughout the month. As a result of the site visitors originated from unusual American residential web service suppliers fairly than from knowledge facilities or outdoors North America, it handed by each of the geography-based and infrastructure-based guidelines Grey had already constructed.
His response added two additional guidelines: extending the prevailing datacenter problem to cowl Azure and different main cloud suppliers, and introducing what he calls a test in opposition to browsers frozen years old-fashioned. Actual customer site visitors confirmed solely 0.54% operating browser variations sufficiently old to set off the brand new rule, with most of that official share remoted to a long-term help launch of Firefox that he exempted from the test.
The revealed rule set and its measured accuracy
Grey’s put up consists of the entire set of 9 Cloudflare Net Software Firewall guidelines presently operating on the location, together with the underlying rule-language expressions in an appendix, protecting nation blocks, named-crawler blocks, a skip rule for Cloudflare-verified bots, continent-level challenges, datacenter-ASN challenges, stale-browser challenges, and a price restrict utilized to extensionless web page requests. The configuration runs on Cloudflare’s Professional plan, which Grey reported prices 25 {dollars} a month.
To check whether or not the extra aggressive problem guidelines have been affecting actual guests, Grey revealed solve-rate knowledge protecting a 48-hour window wherein Cloudflare issued 106,437 challenges, of which 252 have been solved, a price of 0.24%. Nation-level clear up charges in the identical window ranged from a 99.14% failure price in India to a 100% failure price in Iraq, a sample Grey reads as affirmation the challenged site visitors was overwhelmingly automated fairly than composed of inconvenienced official guests.
Grey additionally disclosed value figures for operating the location. Regular month-to-month infrastructure spending sits round 90 {dollars}, operating on Cloudflare Staff with a D1 database and edge KV caching, which means most bot requests hit cached responses at minimal marginal value. Throughout one interval of heavy scraping exercise, that invoice rose by roughly 500%. Grey acknowledged that on a standard digital non-public server billed by CPU and bandwidth consumption, the site visitors quantity he documented would characterize an existential value drawback, whereas on his present edge-caching structure it features as an operational nuisance.
Findings the put up frames as unresolved
Grey’s concluding part reviews that within the 24 hours following implementation of his latest guidelines, Cloudflare blocked 46,729 requests outright, of which 43,150 originated from Amazon’s crawler alone, persevering with to strike the block carried out two days earlier. The identical interval noticed 63,969 challenges issued, of which 552 have been solved.
He frames the underlying drawback as financial fairly than technical, pointing to Cloudflare’s pay-per-crawl framework, below which crawlers pay a price per request on the community edge, as a mechanism he would use have been it commercially accessible to him, stating he can be keen to promote Amazon’s crawler entry to his 3.5 million month-to-month web page views at a good price fairly than blocking it outright. Absent that sort of market, his acknowledged working rule is that any crawler that by no means sends him a customer will get blocked.
Grey’s put up is a single website’s knowledge, not an trade examine, nevertheless it lands inside a physique of measurement that PPC Land has tracked since mid-2024, and the numbers he reviews sit inside ranges prior reporting has already established as directionally constant throughout a lot bigger datasets.
The 35,000 to 1 determine for Claude-SearchBot falls inside the vary Cloudflare itself disclosed when it opened a Bot Management dashboard to clients on July 1, 2026, documenting crawl-to-referral ratios spanning from 118 on the low finish to almost 50,000 on the excessive finish throughout the crawlers it tracks. Cloudflare’s personal historic figures for Anthropic particularly moved from 286,930 crawls per referral in January 2025 right down to 38,000 by July of that 12 months, in keeping with reporting tied to Anthropic’s crawler documentation update revealed February 25, 2026, which clarified the separate roles of ClaudeBot, Claude-Person, and Claude-SearchBot and acknowledged that each one three respect robots.txt. Grey’s put up seems to corroborate that declare instantly, reporting that Claude-SearchBot’s request quantity dropped by roughly 99.96% inside days of a firewall block, in line with a crawler that honors an access-denied response fairly than circumventing it.
The Amazon figures match a equally documented sample. A March 2026 report coated by PPC Land discovered that AI bots crawl retail sites 198 times extra per go to than Google, and famous that Amazon’s Rufus assistant had by that time generated near 12 billion {dollars} in incremental annualized gross sales, a industrial incentive that helps clarify why Amzn-SearchBot’s crawl quantity on Grey’s website practically doubled Claude-SearchBot’s earlier tempo.
The residential proxy sample Grey describes connects to a Samsung security disclosure reported by PPC Land, wherein a researcher discovered proxy code from Shiny Information embedded inside 1 / 4 of sampled Tizen sensible TV functions, changing client units into exit nodes for a similar sort of distributed scraping site visitors.
The broader debate over whether or not blocking AI crawlers helps or harms a writer stays unsettled within the analysis PPC Land has tracked. An April 2026 working paper from researchers at Rutgers Enterprise College and The Wharton College discovered that information publishers who blocked giant language mannequin crawlers by robots.txt misplaced roughly 7% of weekly site visitors inside six weeks, a discovering that stood in some stress with an earlier model of the identical analysis, coated individually, measuring a 23% decline. Grey’s put up doesn’t instantly have interaction that literature, and his website’s site visitors composition, a specialist donor database fairly than a common information writer, might not generalize to the shops that analysis examined.
Cloudflare’s personal July 1, 2026 coverage shift, shifting from charging AI crawlers per particular person fetch towards paying publishers based on whether their content was actually used to generate a solution, was tied by the corporate to inside knowledge exhibiting greater than half of crawl site visitors from bots it classifies as official re-fetches pages unchanged since a earlier go to, a wasted-crawl drawback structurally just like what Grey paperwork on his personal, a lot smaller website. A companion coverage units a September 15, 2026 default below which Coaching and Agent class crawlers will probably be blocked mechanically on advertising-carrying pages for domains newly becoming a member of Cloudflare’s community, whereas Search class crawlers stay allowed by default, a distinction that maps onto Grey’s choice to dam Claude-SearchBot and Amzn-SearchBot whereas persevering with to permit Googlebot and Bingbot.
Grey’s central methodological level, that visitor-tracking instruments constructed on JavaScript execution systematically undercount site visitors really reaching a server, carries direct implications for anybody counting on instruments like Google Analytics or Believable to evaluate infrastructure load or the true value of serving content material to non-paying automated guests.
Timeline
- November 2025 – PatronView information a wave of roughly 4,000 single-page periods with no referrer, later recognized as an early bot sample.
- January 2025 – Cloudflare knowledge exhibits Anthropic’s crawl-to-referral ratio at 286,930 to 1, later cited in PPC Land coverage of Anthropic’s documentation replace.
- July 2025 – Anthropic’s crawl-to-referral ratio falls to 38,000 to 1, in keeping with the identical Cloudflare figures.
- February 25, 2026 – Anthropic clarifies the separate functions of ClaudeBot, Claude-Person, and Claude-SearchBot and commits to respecting robots.txt.
- March 2026 – A report finds AI bots crawl retail sites 198 times extra per go to than Google, with Amazon’s Rufus assistant tied to almost 12 billion {dollars} in incremental annualized gross sales.
- April 21-23, 2026 – PatronView information 3.6 million requests in sooner or later from 361,844 distinctive IP addresses, principally in China; Grey blocks China, then Vietnam and Singapore.
- April 26, 2026 – Up to date Wharton and Rutgers analysis finds publishers blocking AI crawlers lost roughly 7% of weekly site visitors inside six weeks.
- June 2026 – Grey measures Claude-SearchBot’s crawl-to-referral ratio on PatronView at 35,000 to 1 and blocks the crawler on the firewall.
- July 1, 2026 – Cloudflare opens a Bot Administration dashboard documenting crawl ratios up to 50,000 to 1 and individually broadcasts a shift to per-answer publisher payments, alongside a coverage setting a September 15, 2026 default block for Coaching and Agent crawlers on new domains.
- Late July 2026 – Distinctive IP addresses hitting PatronView climb from a baseline of roughly 18,000 to 124,000 on July 31, traced to a residential proxy community.
- August 5, 2026 – Grey disables Cloudflare’s JavaScript Detections function, elevating his cellular Lighthouse rating from 58 to 99 inside an hour.
- August 5-6, 2026 – Grey blocks Amazon’s Amzn-SearchBot crawler, then measured at roughly 117,000 requests per day.
- August 7, 2026 – Grey publishes the complete findings on PatronView and summarizes them on X.
Abstract
Who: Nick Grey, founding father of the philanthropy donor analysis website PatronView, primarily based in Austin, Texas.
What: Grey revealed server log knowledge exhibiting that bots accounted for greater than 99% of site visitors to his 1.5 million web page website, together with a measured 35,000 to 1 crawl-to-referral ratio for Anthropic’s Claude-SearchBot and a 117,000 requests per day price for Amazon’s Amzn-SearchBot, alongside the precise Cloudflare firewall guidelines he inbuilt response.
When: Grey revealed the findings on August 7, 2026, describing occasions throughout roughly a 12 months, from an preliminary bot sample in November 2025 by a firewall change made two days earlier than publication.
The place: The findings concern PatronView’s personal infrastructure, operating on Cloudflare’s community, although the site visitors sources documented span China, Vietnam, Singapore, and distributed residential web connections primarily in the US.
Why: The put up supplies one of many extra granular, independently measured accounts of how AI search crawlers behave on an actual manufacturing website, providing figures that align with broader trade measurements from Cloudflare and third-party researchers whereas giving website operators a selected, documented set of firewall guidelines and their measured results on official site visitors.
Source link

