In a single 24-hour window this year, automated clients hit add-to-cart URLs on WordPress sites 7.67 million times. ClaudeBot accounted for roughly 3.75 million of them. Not one of those requests belonged to a shopper.
That is the shape of what Singapore site owners inherited in 2026. AI crawlers now reach 39% of the top one million websites every month. Cloudflare’s own measurement found that more than half of the requests re-fetch content that has not changed since the crawler’s last visit. A blog post you published in 2023 gets pulled down again this week, and again next week, and the CPU time lands on your hosting bill.
For anyone running a site on shared WordPress hosting, the symptom arrives before the diagnosis. Response times drift. The host sends a resource-limit warning. Analytics show sessions that never convert and never scroll. Teams usually reach for a caching plugin, which helps a little, because the requests are not cached at the edge in the first place and keep arriving at the origin.
Cloudflare moved on this. On 15 September 2026 it replaced the old choice of block-AI-bots-or-do-not with a three-way split. Search covers behaviour that collects or indexes content to answer questions later. Training covers a crawler taking content to train or fine-tune a model. Agent covers automated behaviour acting in real time on a person’s behalf, which is the category that matters most and gets discussed least.
Under the new defaults, Training and Agent are blocked on pages that display advertising, while Search stays allowed. The defaults apply to new domains onboarding to Cloudflare, to new customers, and to existing customers on the Free tier.
Here is the part that most coverage skipped. If you run a B2B site with no advertising on it, the blocking default does not fire. Your pages carry no ads, so there is nothing for the rule to act on. The reason to pay attention is the split itself, not the block.
Why blocking Agent traffic costs you buyers
Think about how a procurement lead now researches a vendor. They ask an assistant to compare three Singapore development partners, and the assistant goes and reads those three sites while the person waits. That fetch is Agent traffic. It is a buyer, mediated by software, reading your capability page.
Block that category and you disappear from the comparison. The assistant returns an answer built from whichever competitors let it in. You will not see a bounce, a failed form, or anything at all in your analytics, because the request never completed. It is the quietest way to lose a shortlist place.
This is the same problem as answer engine optimisation, arriving through a different door. Teams have spent two years making their content citable by ChatGPT, Claude and Perplexity. A blunt crawler block undoes that work at the network layer, below the level where anyone thinks to look.
What maintenance covers now
Website maintenance used to mean plugin updates, backups and an uptime check. Those still matter. Five practices have joined them, and they are all things a retainer should be doing without being asked.
Read the crawler logs every month. Not pageviews. Request counts by user-agent, with origin CPU alongside them. You are looking for a single crawler taking an outsized share, and for repeat fetches of pages that have not changed. Until you have that number, every other decision is a guess.
Set the three categories deliberately, per section of the site. Allow Search and Agent on everything a buyer would read: services, case studies, the blog, contact. Rate-limit or block Training on material you do not want absorbed into a model, such as proprietary frameworks or client deliverables you host publicly. The default is not a policy. It is someone else’s policy applied to you.
Serve a current llms.txt. It gives crawlers a compact map of what is worth reading, which cuts repeat full-site fetches and gives you a say in which pages represent you. It needs regenerating when your content changes, which is why it belongs in maintenance rather than in a launch checklist.
Rate-limit rather than block. Sixty requests per minute per user-agent keeps a crawler from overwhelming a shared host while still letting it read you. Cloudflare Rate Limiting Rules handle this without touching your application.
Watch cache hit ratio and origin CPU, not just traffic volume. Crawler load is cheap when it is served from the edge and expensive when it reaches PHP and the database. Two sites with identical crawler volume can differ tenfold in cost, and the difference is entirely in cache configuration.
The mixed-crawler problem you cannot solve yourself
One complication deserves naming, because it is outside your control. Google, Apple and Microsoft run crawlers that bundle search indexing with training collection in the same client. Block Training from those companies and you may lose search visibility along with it, because the two functions share a user-agent.
Cloudflare’s September deadline was partly an attempt to force those companies to separate the two. Until they do, the honest position is that you cannot cleanly refuse model training by a company whose search traffic you depend on. Anyone who tells you there is a clean setting for this is selling something.
What it actually costs a mid-sized site
Numbers make this concrete. Take a WordPress site with 40,000 human sessions a month, sitting on a mid-tier shared plan, with a cache hit ratio of 70% because the product and search pages are excluded from caching.
Crawler volume on a site that size commonly runs between 150,000 and 400,000 requests a month once three or four AI clients have found it. At a 70% hit ratio, somewhere near 100,000 of those reach PHP. Each one spins up the WordPress bootstrap, queries the database, and returns a page nobody reads. That is the resource-limit email, and it is why the upgrade quote arrives before anyone has diagnosed the cause.
The fix is rarely a bigger plan. Raising the hit ratio to 95% on the same crawler volume cuts origin requests by roughly five sixths, and edge-served requests cost almost nothing. Teams that upgrade hosting first usually end up paying for capacity they then do not need.
One Singapore wrinkle is worth flagging. Sites hosted on local shared infrastructure for data residency reasons often sit on plans with tighter CPU ceilings than the equivalent overseas tier. The same crawler volume that a larger host absorbs without comment will trip a limit here, so the residency decision and the crawler question are connected whether anyone planned it that way.
What to do this month
Pull your crawler numbers first. The conversation changes completely depending on whether AI clients are 3% of your requests or 40%, and most teams have never looked.
If you are on Cloudflare’s Free tier, open the bot management settings and check what the September defaults did to your configuration. If your site carries no advertising, you are probably unaffected, and you should confirm that rather than assume it.
Then decide the Search, Training and Agent question on purpose, written down, per content type. That document is worth more than any plugin, because it is the thing you will hand to whoever maintains the site next.
Webpuppies maintains WordPress and custom platforms for clients across Singapore and the region, including crawler policy, edge configuration and the AEO work that keeps a site citable to AI assistants. If you want your crawler numbers pulled and read, get in touch and we will run the audit.
Sources
- AI Crawlers Are Eating Your Hosting Plan: The 2026 Bot Traffic Problem
- Cloudflare’s new policy pushes AI companies to pay for publishers’ content, TechCrunch
- Cloudflare Blocks AI Crawlers by Default on 15 September 2026, Crawl Lab
- AI Crawler Load on WordPress: Causes and Solutions in 2026
Frequently Asked Questions
What is AI crawler traffic?
Automated requests from clients operated by AI companies, such as GPTBot, ClaudeBot and PerplexityBot. They fetch page content to index it, to train models, or to answer a question a user asked an assistant in real time.
Did Cloudflare block AI crawlers on my site on 15 September 2026?
Only if your site is on the Free tier, is a new domain onboarding, or belongs to a new customer, and only on pages that display ads. Paid customers on existing plans keep their prior settings until they change them.
Should I block AI crawlers entirely?
No. Blocking Search and Agent categories removes your site from AI assistant answers and from research an assistant performs for a buyer. Rate-limit Training traffic instead and keep Search and Agent allowed.
How do I know how much AI crawler traffic my site gets?
Check your CDN or server logs filtered by user-agent, or your Cloudflare dashboard’s crawler view. Compare request volume against your cache hit ratio to see how much is reaching your origin server.
Does an llms.txt file reduce crawler load?
It can. A current llms.txt points crawlers at the pages worth reading in a compact form, which reduces repeat full-site fetches and gives you some say over which content represents you.
