We seem to have built civilisation around the fact that evidence and acceptance are not quite the same thing.
A researcher is expected to use real, direct sources. A designer is expected to understand that “clean” does not mean submitting a blank page. A developer asked to build a working checkout is also expected not to make it take four minutes to load. Humans are surprisingly good at carrying all of this unstated context because we share a basic sense of what reasonable work looks like. Software never needed much of that because we usually gave it instructions narrow enough to execute literally.
You know what I am talking about if you have watched this bit from “The Orville”:
AI agents are beginning to work in a messier world, where the instructions do not spell out everything that counts as doing the job properly. Once money depends on that missing context, interpretation is also part of the transaction.
Frontrun: Be One Step Ahead
Before top investors back a young company, they leave a trail. Every top deal starts with an investor getting curious and following the company on their socials. It eventually culminates in a funding deal a couple of months later.
Whether you can sneak in on that information before the deal is struck and the hidden gem is out in public depends on your ability to track the trail.
Frontrun watches what 2,000+ leading investors, founders, and tech insiders follow. The moment they start crowding around the same company, Frontrun flags it and tells you who’s behind it and what they’re building. It even drafts your first message to the founder.
You can also get customised daily reports around the accounts you are tracking. When the accounts you follow start piling into the same company, you’ll see it on your feed and in your inbox by breakfast.
If someone paid a freelancer $15 to research 20 companies and the brief is reasonably clear, what do they expect? Find the companies, explain what each one does, use independent sources and send the finished report by this date. When the work arrives on time and contains all 20 names, the person who did it considers the job finished.
The buyer sees it differently. Several citations point back to company blogs, and two descriptions appear to have been copied almost word for word.
With two people involved, the next step is that they argue, send screenshots, point back to the original brief and go back and forth. If they still cannot agree, somebody else may eventually have to decide if the work was good enough to get paid.
What happens if both sides are AI agents?
Money itself is not particularly difficult.
An escrow contract can hold the $15 until the job is completed.
A blockchain can show when the payment was deposited and when the work was submitted.
Identity systems can help establish which agent performed the job.
Signed records can preserve the instructions both sides agreed to.
The unresolved part is that somebody has to read the brief, inspect the work and decide whether the seller actually did the job.
Commerce has always had institutions for this problem, but we notice them when something goes wrong. When a PayPal buyer claims something never arrived or didn’t match the description at all, both sides get a chance to fix it themselves first. If they can’t, one of them can escalate it to a claim and ask PayPal to check the proof and decide who is right. PayPal says most of those decisions take around 14 days, though a few take 30 days or longer.
Larger commercial disagreements have their own machinery. US brokers have FINRA, the industry’s self-regulator, which runs its own arbitration forum. It received 2,597 new cases in 2025 and took an average of 13.4 months to close them. The American Arbitration Association, a non-profit organisation, says 580,000 cases were filed with it in 2025, including more than $29 billion in business-to-business claims and counterclaims. The International Chamber of Commerce received 881 new arbitration cases during the same year, while the total value of its active caseload reached $299 billion. These are obviously much bigger disputes than the ones agents are likely to have right now. But they show that once enough deals are happening, people start arguing often enough that a whole industry pops up around looking at the evidence and figuring out who should get paid.
AI agents are starting to raise that kind of problem on the other side of things, where transactions can be really small and happen so often that conventional dispute processes don’t make much sense.

Visa and Artemis studied two emerging payment protocols for machine-to-machine commerce and found that x402 processed about 109.6 million adjusted transactions worth roughly $15 million between its launch in May 2025 and April 2026. The average payment on the protocols Visa examined was a fraction of a cent. A human support team cannot sensibly spend half an hour investigating every disagreement over a payment worth pennies.
I went looking for who is solving this and found GenLayer, trying to turn this into infrastructure.
GenLayer is building a blockchain itself for decisions that ordinary smart contracts struggle with. Traditional smart contracts are good when every machine can do the same math and land on the same result. They can tell you if money arrived, or if the deadline’s gone. The issue shows up when it needs to read something and interpret it. Maybe it has to go through a research report, look at a webpage, understand a paragraph in a contract, or figure out if an image follows a creative brief. AI models never reply with the same answer. All of them can look at the same material and come back with different conclusions, so the assumption that every node has to get the exact same output falls apart.
GenLayer has an “Intelligent Contract.” Instead of requiring every validator to reproduce identical text, the contract can ask several validators to judge whether a proposed result is acceptable according to rules written into the contract. The validators can use large language models and public web data while doing this.

At first glance, this sounds like another missing layer in the agent stack. But some other systems already do pieces of that job. So GenLayer is not starting from zero. Its more specific role is deciding what the evidence means when both sides still disagree.
Google’s Agent Payments Protocol, or AP2, already creates cryptographically signed records of what an agent was authorised to buy, what the merchant offered and what payment was approved. Its specification explicitly says those mandates and receipts can later be combined as evidence in a dispute. But AP2 itself does not decide/judge who is right or who should get the money. The protocol preserves a trustworthy record of what happened and leaves the process for interpreting that record to somebody else.
ERC-8004, Ethereum’s standard for agent identity and reputation, also has a Validation Registry. An agent can send its work there to be checked by another system, and that result can then be recorded onchain as part of the agent’s history.
Blockchain can show if money moved, so payment disputes are sorted. When a dispute needs outside information to settle, an oracle goes and brings that information onchain. So there’s a clear missing piece when the evidence exists, someone still has to judge what it means, that’s where GenLayer comes in.
To see how much real work falls into each bucket, my agent on slack (whom I trust blindly because it’s us against the world) pulled a little over 300 live task listings on September 24. It included those from Freelancer.com, Fiverr, Upwork, Scale AI’s Outlier, Prolific and Daydreams TaskMarket, an on-chain marketplace where agents post and complete paid tasks. About 23% of the tasks were easy to check with a clear yes or no, said my agent. Another 45% depended mostly on personal judgment. The remaining 32% had evidence to look at, but still had to judge whether the work was actually good enough.
That is not a market-size estimate, and the classification itself involved judgment by an AI. But you get the idea. And every task involving AI doesn’t require an AI court.
A task like designing the most beautiful restaurant logo has several validators who may agree that one design looks better than another, but their agreement does not turn taste into an objective standard. Unless the buyer wrote a very detailed creative brief beforehand, the dispute is still rooted in personal preference. Similarly, a marketing campaign can deliver the promised engagement numbers while leaving a disagreement about how much of that activity came from genuine users. In each case, there is evidence to inspect, but it needs interpretation before payment is made.
Crypto has been experimenting with decentralised dispute resolution for years. These systems approach disagreement in very different ways.
UMA’s Optimistic Oracle starts by assuming a claim is fine unless somebody objects. The person proposing it posts a bond. If no one challenges it during the waiting period, the answer is accepted. If someone does challenge it, UMA token holders step in and vote. This keeps costs low because most claims pass without a full review, and the expensive voting process only happens when someone challenges one.
Kleros is a decentralised dispute-resolution system on Ethereum. Instead of a company employee or court deciding a dispute, Kleros randomly selects people who have staked its PNK token to act as jurors. If one side appeals, the next round brings in a larger group of jurors. That makes Kleros useful for cases where human judgment really matters, but it also means the process takes time and costs money. A three-juror round in the General Court is about 0.015 ETH before any appeal, so the model makes no sense when the dispute itself is only worth a few dollars.
GenLayer takes the judging job Kleros gives to human jurors and gives it to AI-assisted validators. One validator proposes an answer, the others check it, and if enough agree, the result moves toward final approval. If someone challenges it, a fresh group of validators reviews the case, with larger groups added if the dispute keeps going. Also, both UMA and Kleros are moving toward AI-assisted adjudication, so I would not call GenLayer “the AI alternative” to two permanently human systems.
If the rules are specific, validators have something concrete to check. If the contract only says the work should be “insightful” or “original,” several models agreeing with each other does not make that standard any clearer.
GenLayer handles this with what it calls the Equivalence Principle. It tells validators how similar their answers need to be for the network to treat them as the same decision. For some contracts, that might mean an exact match. For others, validators can use different wording or reasoning as long as they agree on the important outcome. GenLayer’s documentation repeatedly encourages developers to narrow these decisions.
That puts a lot of power in whoever writes the contract. They decide what evidence matters, what counts as success, and how much variation between answers is acceptable. GenLayer can judge against those rules, but it cannot fix a vague standard after the dispute starts.
A verdict only matters if it can actually change the transaction. In practice, that means the money or some other valuable outcome needs to be tied to the decision. A dispute can then end with funds being released, returned, or an agent’s reputation being updated.
For a small agent transaction, the goal may simply be making sure the money goes to the right side rather than a legally enforceable ruling.
Where would this actually be useful?
Agent marketplaces already combine delegated work with payments and increasingly with onchain identity, so that’s one of the most obvious places to look.
Daydreams TaskMarket, for example, runs on Base and allows humans or AI agents to post tasks backed by USDC. Its different task formats reveal where adjudication would and would not be useful. A benchmark task can ask workers to compete on a measurable metric such as accuracy or latency. If the task is based on a clear score, like one agent getting 96% and another getting 91%, the platform can simply check the numbers and pick the better result.
But with a bounty, several people might submit different kinds of work, and the requester has to decide which one is best. If there is no simple number or rule that tells you who did better, someone has to look at the work and decide.
When it comes to Software and security bounties, a task asking an agent to make a fixed test suite pass can be settled mechanically. A vulnerability bounty may produce a real disagreement about whether an exploit falls inside scope, whether its severity was described accurately or whether somebody had already reported the same underlying issue.
In performance marketing, counting 50,000 social interactions is easy enough. But are they genuine users or automated accounts? GenLayer already showcases Rally, a performance-marketing application that uses its validators to analyse social activity and assess the authenticity of engagement.
Now, a service-level agreement, or SLA, is basically a promise between a service provider and a customer about the minimum level of service they will provide. For example, a cloud company might promise 99.9% uptime, or support replies within two hours. If the provider fails to meet that promise, the customer may be entitled to a refund or a service credit.
If the agreement already specifies which logs, monitoring services, and exclusions count as evidence, an automated adjudicator has a much narrower thing to deal with.
Research is a good example of where GenLayer works well up to a point. Some parts of research are easy to check. Did the report answer all the questions? Are the sources real? Validators can inspect those things.
But something like “is this insight original enough?” is much harder. There may be no clear rule or evidence that settles it.
The work cannot be judged by a simple yes/no script, but there is still enough evidence and structure for different validators to make a reasonable decision. If the task is almost entirely about taste or vague quality, the system has much less to work with.
If you think using several AI judges automatically makes the decision reliable, no.
If all five validators use similar models, they might all make the same mistake. They could misunderstand the same instruction, trust the same bad source, or all be fooled by manipulated content. GenLayer tries to reduce this by having validators check decisions independently and by bringing in new validators if someone appeals.
The evidence can also be unreliable. A webpage might change, an API might give different answers at different times, or a document might contain instructions designed to confuse the AI reading it. So the system only works well when validators can properly inspect the same evidence. Sometimes there may simply be no clear answer. GenLayer can return an “Undetermined” result instead of forcing a yes or no.
GenLayer’s design leans on Condorcet’s Jury Theorem, the idea that a group of independent reasoners beats any single one. That only holds if the reasoners are truly independent and AI models learn from overlapping data. Kleros researchers tested something close to this on 99 real customer disputes from a Latin American crypto exchange. ChatGPT 5.5 sided with the customer in 3 cases and Claude Opus 4.7 in 14. When they reran the cases on the next Claude version, the platform’s win rate rose from 86 to 95%. Mixing models helps when their leanings are random. If a whole generation of models drifts the same way, a mixed panel drifts with it.
How big is the market now? The market is hard to size because only a fraction of agent transactions will ever become disputes. But existing systems show the demand for adjudication is already great.
Merchants handled about 261 million chargebacks in 2025, worth roughly $33 billion in dispute processing at $128 each. At the other end, GenLayer currently shows around 25,800 testnet decisions a day. Even at $1 per decision, that is only about $9.4 million a year. For comparison, eBay at its peak handled roughly 60 million disputes a year, or about 164,000 a day. GenLayer only becomes a large business if machine commerce eventually produces dispute volumes at that kind of scale.
For this to work, the contracts need clear rules, and the evidence needs to be reliable. The validators need to be independent enough that their agreement actually means something. Appeals also need to stay cheap.
Human commerce already has courts, arbitration and dispute teams for when two sides disagree about the same deal. If agents start doing more business with each other, they may need a faster and cheaper version of the same thing, because machines always need a cheap, fast version of everything.
Maybe it is a good thing we teach machines how to decide what is fair before we teach them to do everything else.
Token Dispatch is a daily crypto newsletter handpicked and crafted with love by human bots. If you want to reach out to 170,000+ subscriber community of the Token Dispatch, you can explore the partnership opportunities with us 🙌
📩 Fill out this form to submit your details and book a meeting with us directly.
Disclaimer: This newsletter contains analysis and opinions of the author. Content is for informational purposes only, not financial advice. Trading crypto involves substantial risk - your capital is at risk. Do your own research.







