From the Crawl to the Citation: What Actually Changed in Agent-Internet Monetization

In twelve months Cloudflare changed its mind about how to charge. The full map of who’s building the toll booth — and which businesses survive the next turn.

On July 1, 2025, Cloudflare announced that AI companies would have to pay every time a crawler fetched a page. Exactly twelve months later, on July 1, 2026, the same company declared that model insufficient and proposed a different one: pay publishers only when their content actually shows up inside an answer.

The distance between those two dates is the best available measure of how fast the economics of machine-read content are moving. It also carries an uncomfortable lesson for anyone thinking about building here: the infrastructure is changing faster than the theses being built on top of it.

It’s worth opening the full map, because there are good businesses hiding underneath a narrative that has already aged twice — and most of them aren’t where the headlines point.

The bargain that came apart

For decades the trade was simple and nobody had to sign it. You let search engines crawl your site; search engines sent you people; you monetized that visit with ads, email capture, a subscription, or an affiliate link. The crawler got the content, you got the visitor.

The 2026 numbers show that bargain coming apart almost brutally. AI crawlers accounted for 52% of all crawler requests in June 2026 — up from 22% in spring 2025. And the ratio between pages crawled and visitors returned, for the leading bots, runs from 118:1 to nearly 50,000:1. On the user side, zero-click search climbed from 56% to 69% in a single year after AI Overviews, and 35% of US consumers now begin product discovery inside an AI tool, against 13.6% who still start with traditional search.

Translated: the content still creates value, but the site stopped receiving its side of the deal. That part of the diagnosis is no longer seriously disputed.

What Cloudflare is building, minus the marketing

Cloudflare’s answer has evolved in three moves.

The first was visibility and control: knowing which AI crawlers access what, and being able to allow or block each one. The second was pay per crawl, which put HTTP status 402 — reserved and ignored since the 1990s — to work as a per-request charge. The crawler requests a page, receives a 402 Payment Required with the price in a header, and retries agreeing to pay. Cloudflare verifies at the edge, before the request reaches the origin, and acts as merchant of record.

The third move is the one almost nobody noticed. In July 2026 Cloudflare began sorting crawlers into three buckets — search, agent, and training — and blocking the agent and training categories by default on any page carrying ads. The change applies to new domains and to the entire free tier starting September 15, 2026. Alongside it came a pay-per-use model: the publisher is compensated when their content feeds an answer, not merely when it gets fetched.

That’s a meaningful course correction, and it says something about the original design: charging per crawl paid badly, because most crawls generate no value at all. The unit of account moved from the crawl to the citation — from being fetched to being used.

But it isn’t just Cloudflare — and the toll booth is already a shelf item

Focusing on Cloudflare to understand this shift is like having watched only Netscape in 1996. It’s the most visible piece, not the most important one, and certainly not where most of the money will end up.

The first thing to accept is that charging crawlers has stopped being a competitive differentiator. The RSL standard — Really Simple Licensing — has been adopted by Cloudflare, Akamai, and Fastly, backed by Yahoo, Ziff Davis, and O’Reilly Media, with Creative Commons joining through its CC signals project. The three largest edge infrastructures in the world now speak the same licensing language. TollBit, for its part, is already integrated with Akamai, WordPress VIP, and DataDome, letting millions of sites switch on monetization with one click instead of building their own payment system.

When three competitors adopt the same standard and ship the feature free in the base plan, that feature is a commodity. If your business idea was “build a pay-per-crawl,” it died before it launched. The opportunity sits above or below that layer, never in it.

Two accounting philosophies — and the choice matters

The next layer up, the licensing intermediaries, has split into two philosophies that produce meaningfully different financial outcomes.

On one side, charging per fetch. TollBit is the example: every URL fetch generates a fixed payment to the publisher at a publisher-set rate, regardless of whether the resulting AI answer cited the source meaningfully. The publisher keeps 100% of the revenue; TollBit charges the AI buyer a transaction fee. The company raised a $24M Series A led by Lightspeed and has expanded into Japan through local media houses.

On the other, charging per citation. ProRata, founded by Bill Gross, attempts to calculate how much a specific answer relied on each article and pays proportionally to that attribution, splitting revenue 50/50. More than 500 publishers have signed up.

The choice between the two isn’t ideological, it’s arithmetic. For commoditized content — general news, B2B guides, encyclopedic material — per-fetch pays more, precisely because that kind of content scores low on per-citation contribution. For differentiated content — primary research, exclusive financial data, hyperlocal coverage that exists nowhere else — attribution pays considerably more per event.

There’s a wide dozen of companies in this space beyond those two: Sphere AI, ScalePost, Created by Humans, Miso.ai. Brookings mapped the sector going from a handful of startups in 2024 to more than a dozen today, and issues a warning worth heeding: nearly all are venture-backed and therefore exposed to acquisition by exactly the Big Tech firms they nominally protect publishers from.

The platform eating the intermediary

That warning is already materializing from above. Microsoft launched the Publisher Content Marketplace in February 2026, co-designed with the Associated Press, Vox Media, Condé Nast, Business Insider, and People. It works like an app store for licensing: the publisher registers the archive, sets terms and pricing, and gets paid when the content is used to ground an answer — starting with Microsoft’s own Copilot, with usage reporting flowing back to the publisher.

What Microsoft did was absorb, with its own scale and distribution, exactly the function a dozen startups were trying to occupy. And it isn’t alone: Perplexity runs a revenue-share program, and Meta has struck direct licensing deals.

The lesson for builders: don’t build anything that depends on being the generic intermediary between publishers and AI labs. Three trillion-dollar companies are entering that space simultaneously.

The rails exist. The demand doesn’t.

The money layer is the loudest and the least settled.

x402 became a foundation under the Linux Foundation in April 2026, with Cloudflare, Coinbase, Google, Visa, AWS, Stripe, Circle, Anthropic, and Vercel among its members — stablecoin settlement, designed for per-request micropayments, already past 100 million cumulative transactions on Base. In parallel, Stripe opened its Machine Payments Protocol to developers in March 2026 as the fiat alternative, using session-based aggregated billing rather than per-call micro-settlement. The same Stripe, with OpenAI, maintains the Agentic Commerce Protocol, which already runs inside ChatGPT and has been adopted by PayPal, Salesforce, and Shopify. Google has UCP with Shopify for the full commerce journey, plus AP2 for authorization. Visa and Mastercard occupy the agent identity layer. AWS decided not to choose: Bedrock AgentCore integrates x402 natively, with reference implementations on CloudFront and Lambda@Edge for the merchant side.

On paper it’s the dream: payment embedded in the HTTP request itself — no checkout, no account, no procurement. The volume tells a different story. As of March 2026 x402 was moving roughly $28,000 a day, and a meaningful share of that was testing or speculative activity, including a pay-to-mint experiment that by itself produced hundreds of thousands of repeated transactions. More telling still: transactions above $1 jumped from 49% to 95% of volume, while the 10-cent-to-a-dollar band collapsed from 46% to 4%.

The economic weight is moving in the opposite direction from micropayments. The promise of “a fraction of a penny the agent won’t even notice” is precisely the segment that’s shrinking.

Four standards coexisting at different layers means nobody knows where this lands. The practical read is twofold: treat the rail as a plugin, not as architecture, and don’t build anything whose viability depends on per-request micropayments working in 2026. Anyone who wires a product to one specific protocol now will be rewriting it in 2027.

The extension almost nobody is watching: tools, not content

The entire public conversation about agent monetization revolves around content — publishers, articles, copyright. That’s the part with the most lawyers and the least opportunity.

The interesting part is the other one: agents don’t just consume text, they call tools. And that layer has striking numbers on one side and a void on the other. The Model Context Protocol SDK went from roughly 2 million monthly downloads at launch in November 2024 to more than 97 million per month by early 2026. Every relevant assistant — Claude, ChatGPT, Copilot, Gemini, Cursor — speaks MCP natively. There are more than 12,000 published MCP servers.

Fewer than 5% of them earn a dollar.

The billing infrastructure already exists and is mature: Apify hosts MCP servers on a pay-per-event model with an 80% share to the developer, retiring flat monthly rental in October 2026; MCPize passes through 85%; Glama, Smithery, and other directories compete on discovery; Nevermined offers outcome-based billing, where you charge per successful result rather than per request — the model that makes most sense when the value is the answer, not the call volume.

Community-reported revenue figures are modest and lopsided: there are accounts of a solo MCP server reaching $10K in monthly recurring revenue within weeks, a middle band in the low thousands per month, and an overwhelming majority at zero. The failure patterns are consistent: too generic a tool, billing bolted on afterward instead of designed in, no distribution, maintenance that stops.

Even so, it’s a better opportunity than content for three reasons. There’s no copyright litigation hanging over it. Demand is verifiable before you build — you can measure whether agents already attempt that job and fail. And differentiation is technical and defensible, not legal.

The trust layer nobody wants to build

One piece is missing from this whole arrangement, and it’s the least glamorous: who audits?

Every pay-per-use or pay-per-citation model depends on the buyer honestly reporting how much they used. The publisher receives a usage report from the same company that owes them money. In practice, the seller is trusting the buyer’s accounting — a situation that in any other market spawns an entire independent audit industry.

The identity layer is getting built: Visa launched its Trusted Agent Protocol with Cloudflare in October 2025, signing agent identity into HTTP headers for merchants to verify against Visa’s directory. But identity answers “who are you,” not “did you accurately declare what you consumed.” Independent usage verification is a business that doesn’t properly exist yet and that every mature market eventually acquires.

What actually holds up as a business

Three businesses circulate in conversations about this shift. They are not equal in quality, and it’s worth ranking them by durability rather than by excitement.

1. The niche data refinery

Pick a niche where the valuable information is scattered, changes constantly, and is annoying to collect. Turn it into clean, structured data. The raw material already exists — maps, reviews, job posts, permits, pricing pages, PDFs — and the entire job is the refining.

The canonical example is med spas: the owner wants to know what competitors charge, what reviews complain about, who’s hiring, which treatment started being promoted in the last sixty days. With that data cleaned up, an agent delivers real recommendations instead of generic prose. But it works identically for roofing, real estate investing, e-commerce, or law firms. The filter is always the same: the data must be valuable, recurring, changing, fragmented, and annoying to obtain — the margin lives in “annoying.”

The wedge is deliberately small: one niche, one city, a hundred businesses, a manual spreadsheet. And the first customer is almost never the end business owner — it’s whoever already sells into that niche. Agencies, consultants, and implementation shops buy market intelligence because it improves work they already deliver. Selling “agent-readable competitive analysis” to a clinic owner means selling a phrase they don’t understand.

This is the strongest of the three precisely because it doesn’t depend on the agent thesis to exist. It’s market intelligence sold as a service, which has always worked. The agent layer is upside, not premise.

The real risks, which rarely come up: privacy law when collecting reviews and identifiable personal data, source terms of service, the ongoing cost of maintaining freshness, and the fact that agents themselves are getting competent at collection. Your moat isn’t the data — it’s reliability, latency, predictable cost, and legal cover. If you can’t articulate why buying from you beats pointing an agent at the same sources, there’s no business.

2. Agent readiness

The idea is helping companies become easy for AI systems to understand, trust, compare, and recommend. The wedge is elegant: run thirty or fifty buyer-intent prompts across the major tools, show the founder what AI says about their company today, and sell the fix. You’re not selling the future — you’re selling the screenshot.

It works. It generates cash in month one. And it’s the most perishable of the three.

By 2026, generative engine optimization is an established software category with more than a dozen consolidated platforms — Profound, Otterly, Peec, Scrunch, plus modules inside Semrush and Ahrefs — in a US market estimated around $365 million. The screenshot you sell as diagnosis costs twenty dollars a month in any of them.

And one data point hits the proposed fix directly: 82% of AI citations come from earned media, not owned or paid content. Cleaning up llms.txt, restructuring documentation, and making the pricing page parseable helps — but it’s a minority of the problem. The larger lever is being mentioned by third parties, which looks far more like PR and community presence than site architecture.

Then there’s the measurement problem: models are non-deterministic, personalized, and change versions without warning. A good share of your monthly “improvement” report is noise. This is sellable, but it’s honest to treat it as vertical consulting, not a defensible SaaS.

3. Turning an expert archive into a tool

Take years of an expert’s content — videos, podcasts, newsletters — and package it into a tool that does one specific job. Not “chat with the expert,” which is too broad and produces a search box with confidence. Something like: paste your cold email, the agent critiques it using that expert’s principles, cites the source lessons, rewrites it, and suggests one test.

It’s the most fun and the weakest as a business. The creator owns the distribution, the brand, and the trust — so the creator captures the surplus. You negotiate rights, build the tagging structure, deliver, and keep a service margin. It’s a good agency. It isn’t a startup, unless you can assemble the same asset across many experts and become the default layer between them and the agents.

Five readings for anyone building here

The bottleneck is always demand, never infrastructure

One publisher running TollBit reported zero revenue: monitoring works, but monetizing requires AI buyers willing to pay, and that market hasn’t materialized at scale. On the tool side, 95% of MCP servers earn nothing. On the rail side, x402 moves less per day than a neighborhood bakery. The classic mistake here is adding more supply to a market already suffering from a surplus of it.

The infrastructure layer is closed

Cloudflare, Akamai, Fastly, Microsoft, Google, Stripe, AWS, and Visa have already taken the toll booth, the marketplace, the rail, and identity. Competing there is irrational. Building on top is rational.

The long tail has nobody serving it

Every intermediary ran toward the large publishers — AP, Condé Nast, Vox. The millions of small sites, sector databases, trade associations, and local data producers have no representation, no bargaining power, and no idea a market exists. Aggregating that long tail by vertical is a real business.

Outside the US it’s open territory

None of these intermediaries operate seriously in Latin America, and the same is true of most of the world outside the Anglosphere. Publishers, sector databases, and holders of structured public data in those markets have no monitoring, let alone monetization — plus a local privacy-law layer none of these players have modeled. Whoever builds it locally starts without competition.

Tools beat content

Content has oversupply, litigation, and reluctant buyers. Tools have verifiable demand, technical differentiation, and no lawyers in the way.

What to do with this

The most common mistake here is confusing an infrastructure announcement with the opportunity. Cloudflare is building the access and payment layer, and has already redesigned it once in twelve months. If your business plan requires that layer to work a specific way, your plan has an expiration date.

What doesn’t expire is the scarcity on the other side: AI systems need clean, trusted, current resources to do good work, and the internet remains a messy warehouse. Whoever organizes one small, specific slice of that warehouse has something to sell today, manually, to human customers — and owns the asset if and when per-request pricing matures.

Five questions help you find your slice:

  1. What decision is expensive?
  2. What information is messy?
  3. What changes often?
  4. Who already pays for help with it?
  5. What would an agent need to do the job better?

The agent internet has no owner and won’t get one. It has seven layers, four competing standards, and a market where almost everyone is building the wrong side of the equation. The question that matters — the one that survives any infrastructure announcement — remains the same: is there someone willing to pay for this today, with money already in their budget? If the answer depends on the market maturing, it isn’t a business yet. It’s a bet.