For most of the web’s life, search gave you a list and left the work to you. We have headlines in blue and hyperlinked.
Then we clicked on whatever felt close to the query. But explaining this as history feels weird and also makes me feel a little old.
Always hated the concept of SEO as a writer. Keywords ruined some good parts of the story you wrote, and citations were counted as a vote when it comes to google ranking. Essentially, your page mattered more if other pages pointed to it. But fine, no complaints. Some kind of credibility is better than being credible. Ads pay for everything on the internet. But publishers hated this setup for obvious reasons.
This is fading away now. I want to know where to find a book on finance and the economy, or what to do when my ankles hurt after a long walk. I go straight to gemini. Some go to perplexity, some to gpt. Some who care less about tokens go straight to claude. Are all of these options better than reading a whole article? We don’t know that… probably not, but it sure makes existence easy. You don’t have to read 2 articles that say two different things. AI has it figured out.
I think those under 30 at least have stopped thinking of search as a place you go.
Then came the agents. They broke the arrangement properly. Non-human traffic overtakes human traffic. An agent running a research job might touch a thousand pages and produce one memo, unlike us, so ads could be of no value in the agentic economy. I wrote two pieces on whether software can pay for what it takes. We concluded that’s the easy part.
An agent navigating the web to complete a task and settle value requires a discrete six-layer stack.
An agent that has to find a page, trust the result, trust the other bot, confirm the work, pay, and leave the writer something. Today we walk through those layers in that order.
Crypto Investing On Autopilot
You’ve probably had your brain fried tracking 50 things for the few tokens you handpicked from a pile of random. Crypto is not a joke. If anything, it is a sad place that robs people off their money. Glider builds on the opposite philosophy. It helps you pick a basket of assets and programmatically rotate it based on conditions you define.
One basket holds the Mag7 at equal weight.
Another mirrors congressional stock trades, the Pelosi Tracker.
Everything stays non-custodial. No gas fees either.
Today’s issue was sponsored by them, and it comes in the week they launched ATPs with Bitwise. More on their release on Bloomberg here. Explore how Glider combines RWAs and programmatic finance directly in its product here.
Let your strategy run itself at Glider.fi.
What is on the live web?
Now that machines and humans both share the word “search”, let’s look at what each wants from a page.
Human search lines up whole pages, so you can scan through titles and decide where to go. Machine search is for AI Agents. An agent doesn’t browse because it has a context window, which is a short-term memory. So instead of returning all the links in optimised order, it pulls the few paragraphs most useful for the task and feeds those in, like a ready-to-use chunk of text. So the memory doesn’t fill up. Also, it just takes milliseconds.
Parallel and Exa are the major players there. Both are independent companies that cater to agentic search. Parallel is Parag Agarwal’s company, the one he started after twitter. It raised $100 million in April at a $2 billion valuation. Parallel sells search against a stated objective and returns compressed excerpts. These are ranked based on how much they help the task, and they provide per-fact citations and confidence scores.
Coming to Exa, it raised $250 million in May at a $2.2 billion valuation, led by a16z. It claims to power Cursor, Cognition, hubspot, OpenRouter and Monday.com across more than 5,000 companies and 40k developers. Exa focuses on meaning instead of keywords. If you search for something and it returns the right pieces even when they never use those words, and it hands back clean page text in the same response.
According to Parallel’s docs, search costs $1 per thousand requests on its two fast tiers and $5 per thousand on its two slower ones. Extraction is $1 per thousand URLs, flat, no matter how long the page is. For Exa, search cost is $7 per thousand requests, 7x what Parallel charges. Exa’s page contents cost $1 per thousand pages, and an answer with citations should cost $5 per thousand.
Was the search any good?
Google usually judges itself in private. Does it do so by checking whether people clicked the first result, stayed on the page or bounced? They also have paid human raters to grade results. There were tests like RTEC, but again, there wasn’t a public leaderboard you could see or buy from. Google Analytics and Similarweb measure traffic, not search quality.
In the agentic internet, there are lab benchmarks that do this. These searches determine whether an AI using that search tool gets the answers right.
OpenAI released BrowseComp in April 2025, which has 1,266 questions built backwards. The writers started with a known, checkable fact, then wrote a question that hides that fact behind several constraints. If you write the question first, like a trivia quiz, models often get it from memory or one lucky search. Writing the fact first, then hiding it, is how you test whether the tool can actually look things up.
Plain GPT-4o scored 0.6%. GPT-4o browsing scored 1.9%. OpenAI’s Deep Research got 51.5%.
The goal is to see whether the system can find one hidden, verifiable fact on the live web and write it down. Each of the 1,266 items has a short official answer. The model gets one try. If its answer matches, it scores a point. If it does not, it scores zero. The percentage is the number of points divided by the number of questions. Plain GPT-4o at 0.6% means it almost never knew the fact already. GPT-4o with browsing at 1.9% means a basic search tool barely helped. Deep Research at 51.5% means a longer search-and-read loop found the fact about half the time.
Like BrowseComp, there are other tests such as SimpleQA (OpenAI; 4,326 questions), an easy-to-grade factuality test. BrowseComp was built because SimpleQA got too easy once models had a fast browse tool. Then there was GAIA (Meta / Hugging Face; 466 human-written tasks), real-world assistant questions that may require web browsing, plus other tools, files, and multi-step reasoning, not search alone. DeepSearchQA is the multi-search research set that Artificial Analysis averages over.
FRAMES (Google DeepMind; 824 items) has multi-hop questions that require pulling facts from several sources and combining them.
Artificial Analysis, a private firm co-founded by George Cameron and Micah Hill-Smith, launched a Search Index on 18 August 2026. In that test, the AI that writes the answer stays the same, and the rules of the test stay the same. The only thing they change is which search API the agent is allowed to call. If the score moves, it is because of the search tool, not because a smarter model showed up.
With no search at all, the same model scores 33 on the Index and 17 on the BrowseComp slice. Those baseline numbers have not moved, even though the search companies’ rankings have changed.
On Artificial Analysis’s Search Index as of 27 August, Perplexity’s mid-tier leads at 80. Parallel’s advanced tier and Brave tie at 75. You.com and Exa sit at 74. Firecrawl, Parallel basic, and Parallel fast all land on 73.
If a search tool beats the baseline, it is actually helping the model find answers. I.e, first, the AI answers from memory. Then it can search. If the score rises, the search has found something new. If it does not, the search did nothing to help the model.
However, the tests are limited because the questions and the model are not yours. Vendors can aim at a known exam, such as DeepSearchQA, the same way sites once aimed at Google.
Who is the other bot?
When two AI agents hire each other, a wallet and a bio aren’t enough to establish trust. A wallet only proves they have funds, and descriptions can easily be faked, leaving no proof of who built the bot or if its past work is real. ERC-8004 is a standard designed to solve this by verifying an agent’s true identity and reputation.
ERC-8004 is the standard built for this. Three registries have been live on the Ethereum mainnet since 29 January 2026, written by people from MetaMask, the Ethereum Foundation, Google and Coinbase.
First, the Identity Registry is an NFT. Each agent gets a numbered token whose URI (Uniform Resource Identifier) points to a registration file listing its name, endpoints, whether it accepts x402 payments, and which trust models it supports. Whoever holds the token owns it, and if it changes hands, the old payment wallet is wiped, so the new owner must sign again to prove they control it.
Second, the Reputation Registry accepts signed feedback. This is a public review system. After you use an agent, you can leave a review on-chain - a score plus labels for what the job was (for example, search or grading). The agent’s owner cannot review themselves. A review can be withdrawn. Anyone can attach a short note under it, so a refund or a spam warning can appear next to the score. It is a public comments thread tied to that agent’s ID, not a private star rating inside one app.
Finally, the Validation Registry is the hook for independent checking. An agent requests validation, and a validator smart contract responds with a score from 0 to 100. It also comes with a link to proof, whether that evidence came from re-running the job, a zero-knowledge proof, or a secure chip attesting to what it saw.
The spec does not handle payments. It says that in so many words. It only records who an agent is, what people said about it, and whether someone checked its work. Payments are someone else’s job. The registry is also thin. By July, 8004scan listed 385,998 registered agents across 29 chains and more than 460,000 reviews. Most of those IDs are empty badges. 89.2% do not publish a standard way to call the agent. Only 8,631 had at least one working service, which is 2.24% of the list.
A May crawl of Ethereum, BNB Chain and Base found that on-chain review scores cannot be compared across agents. Because most feedback isn’t tied to a job anyone can check, and fake reviewers are cheap. After the flagged reviews were removed, a large share of rated agents had nothing usable left. The standard publishes those signals in one format and leaves someone else to decide which reviews are real. ENS, EigenLayer, The Graph and Taiko have said they will plug in, but they haven’t built that filter.
Did agents do the work?
Agentic Commerce architecture (co-designed by the Ethereum Foundation and Virtuals Protocol) is built to let autonomous AI agents hire one another, coordinate work, and settle payments on-chain without human middlemen. ERC-8004 and ERC-8183 (the escrow standard) come with that. The latter standardises how tasks are posted, funded, verified, and settled between agents.
How does it work? The client agent posts a job specifying the task description, deadline, budget, and a designated Evaluator. The Client deposits funds into the smart contract. The money is now locked, and neither party can access it unilaterally.
The Provider Agent runs the compute off-chain. Then it uploads the final result to a decentralised storage network (like IPFS or Arweave) and submits only the cryptographic hash or URI on-chain. The Evaluator (which can be a specialised judge agent, an automated oracle, or a zero-knowledge verifier) inspects the deliverable.
The evaluator either completes the job, paying the provider minus fees, or rejects it and refunds the client. If the Provider disappears or the Evaluator stalls past the expiry deadline, a public function unlocks and refunds the Client to ensure that money cannot remain permanently trapped.
The evaluator can be the client too. The spec offers that as the normal setup when there is no third party. And it states, “No dispute resolution or arbitration; reject/expire is final.” If the person or bot you pick to grade the work decides to rip you off, there is no manager or customer support to appeal to. You are on your own, and you can’t even rely on their online ratings to pick a trustworthy one because most of the reviews are fake, as we discussed in the last part.
While systems like UMA, Kleros, and Bittensor mitigate single-point-of-failure evaluators, they trade off speed and cost. A purely automated escrow settlement settles in seconds for fractions of a cent, whereas dispute escalation through UMA or Kleros requires challenge windows, staked bonds, and human-in-the-loop voting periods.
Did they move money?
We have been talking about this a lot lately. So I will be quick.
x402 is Coinbase’s creation. It is a layer on top of existing chains that charges per request and moves stablecoins.
Read: $5 Spent Well
Then there is the Machine Payments Protocol, published in March 2026 and co-authored by Stripe and Tempo, the same day Tempo’s chain went to mainnet. MPP is an HTTP standard. The server returns a 402 Payment Required with a challenge. The agent retries with a payment credential. The server then replies with the resource and receipt.
How the money actually moves depends on the rail. On Tempo, a one-off call can settle on-chain in about half a second. For a string of small charges, the agent can lock some money first, then sign a running tab as it goes. The service cashes that tab on the chain at the end, not on every click. The cap is whatever was locked in that session or allowed by the access key, not a universal “approve once, spend forever” switch.
MPP can use several rails. Visa published a card specification so agents can pay with tokenised cards. Lightspark has a separate Lightning charge method for the same 402 flow and a separate Visa card programme. A merchant who already takes card payments through Stripe can take an MPP card payment into a normal Stripe balance. Lightning is a separate option under the same MPP rules. Stripe also supports x402. Both MPP and x402 use the HTTP 402 “please pay” step. They are two different protocols.
Did the writer get paid?
If they don’t, then nothing is left to search. If writers do not make money, they will stop publishing, and AI agents will run out of new material to read.
Cloudflare, Parallel, and OpenLedger are three well-known names here.
Read: Paying Is Easy
Cloudflare originally charged AI companies a flat fee every time a bot scraped a webpage. That was too basic because scraping a homepage cost the same as scraping an in-depth investigation. Now, Cloudflare is switching to pay-per-use, a system that pays publishers only when their content is actually used to create value. They are also starting to block AI crawlers on ad-supported sites unless the AI company agrees to pay.
Parallel launched Index on 19 May. It uses a Shapley calculation, borrowed from cooperative game theory, to work out how much each source contributed to a finished agent task and pays accordingly. Instead of paying a flat rate per click, it uses a math model to measure how much a specific article actually helped the agent complete its task. Launch partners include The Atlantic, Fortune, PR Newswire, PitchBook, ZoomInfo, Tracxn, RocketReach, Enigma and Fiscal AI.
Right now, Index only pays out when agents use Parallel’s own search tools. Until other AI platforms adopt it, it can’t be a universal standard for the entire web.
OpenLedger is another version of this, but it uses tokens. It tracks how much a dataset influenced an AI model and pays contributors in its own token. It was designed to reward developers for AI training data rather than compensate journalists for daily articles.
I went looking for what any of these have actually paid a working writer. Cloudflare hasn’t published payout figures. Parallel hasn’t published payout figures. OpenLedger pays in OPEN and publishes no independent totals. So we have no receipt to come to a conclusion here yet.
When humans built trade, the cash register was the last thing to arrive. We took years to figure out handshakes and local reputation before we invented frictionless payments. The machine web seems to be doing it backwards. It built instant, streaming global settlement on day one. The agentic economy is now frantically trying to invent an on-chain small-claims court so two bots don’t rob each other over a few-dollar task.
I’ll leave you with an observation from Terry Pratchett.
“It was all very well to go on about pure logic and how the machine was only doing what it was told, but people had a right to expect machines to have a bit of common sense.”
That’s it for the day.
Token Dispatch is a daily crypto newsletter handpicked and crafted with love by human bots. If you want to reach out to 170,000+ subscriber community of the Token Dispatch, you can explore the partnership opportunities with us 🙌
📩 Fill out this form to submit your details and book a meeting with us directly.
Disclaimer: This newsletter contains analysis and opinions of the author. Content is for informational purposes only, not financial advice. Trading crypto involves substantial risk - your capital is at risk. Do your own research.










