If I call a plumber, I pay when the leak stops. If a company hires ten scientists to build a better battery, it pays them every month whether or not the battery ever shows up.
It is what you do when you cannot tell in advance if an answer exists, and when the only people able to judge the work are others who do the same work.
The oldest alternative is the prize. In 1714 the British Parliament offered up to £20,000 to anyone who could find a ship’s longitude at sea. John Harrison, a carpenter who taught himself clockmaking, spent about forty years building clocks for it. His fourth one kept time well enough on a sea trial that ended in 1762. He then spent another decade arguing with the Board of Longitude over whether the trial counted, and got most of his money in 1773, after King George III stepped in. The prize drew in a self-taught clockmaker from outside the scientific establishment. It also became a very long fight about who gets to say the test was passed.
Prizes never became the main way to fund science. Universities and, from around 1900, company labs put researchers on a salary. After the Second World War, governments began funding the research no company would, and the grant became the basic unit of academic science. A scientist writes a proposal, other scientists score it, and the money arrives before the work begins.
A salary buys a person’s time whereas a grant buys a plan. The US National Institutes of Health received 62,592 applications for research project grants in fiscal 2025 and funded 13% of them, at an average of about $675,000. Reviewers judge the proposal and the record of whoever wrote it, because there is nothing else to judge yet. Companies do the same with salaried teams. WIPO’s Global Innovation Index 2026, released 29 September 2026, says corporate R&D among the world’s largest spenders hit a record $1.5 trillion in 2025, up close to 8% nominally and 5.8% in real terms, and equal to 6% of those firms’ revenue, the highest share on record.
When a company doesn’t want the headcount, it rents it. Contract research firms such as IQVIA, which reported $16.3 billion in revenue in 2025, run lab work and clinical trials for drug companies for a fee. They get paid the same if the drug fails.
The exception has always been the kind of problem where an answer is cheap to check. In 2006, Netflix published a pile of movie ratings and offered $1 million to anyone who could beat its recommendation system by 10%. Kaggle turned the Netflix Prize into a service companies could buy.
The contestants in that kind of scored contest changed now. AI coding agents can now run the whole loop on their own. They read the code, try a change, run the test, keep the change if the score got better and throw it away if it didn’t.
Yukon is a bet on that loop, made by a crypto company. It launched on August 12, so there are now seven weeks of results to look at. Nothing stops us from doing that. Let’s take a closer look.
Frontrun: Early Bird Holds the Edge
Most people hear about a promising startup the day it announces its funding. By then, the investors who backed it have known about it for months. How did they get there first?
Hours of digging, who started the company, what they’ve built before, how to reach them, and understanding what to write to them so that you can join them early in their journey.
Frontrun does this digging for you. When top investors start following a young company on X, Frontrun flags it, finds the founder, and writes your opening message. You can read the draft, tweak it, and hit send without leaving the platform.
Frontrun gives you a ready brief on every company it flags.
You can also automate a daily morning report with the list of startups to track based on what your favourite accounts are following. Served hot, along with your breakfast.
Yukon comes from Eigen Labs, the Seattle company behind EigenLayer. It started with a paper Google chose not to publish in full.
In March, Google’s quantum team said it had found a much cheaper way for a future quantum computer to break the signatures that protect Bitcoin and Ethereum wallets. It judged the recipe too risky to release, so it published a cryptographic proof that the recipe exists, along with a program that could check any such recipe and count its cost.
Gautham Anant, a 22-year-old engineer at Eigen Labs, was taking an introductory quantum computing course at the time. He realised that software could check whether a proposed attack would work and give it a score based on how much computing power it needed. That made it possible to compare different designs, with a lower score meaning a less demanding attack. Google’s published result scored about 3 billion on that scale. Anant put AI agents to work trying to bring the score down. They improved their early attempts quickly, but progress stalled before they could match Google.
On June 1, Eigen opened it to anyone, under the name ECDSA.fail. The score combines how much quantum memory a design needs with how many costly computing operations it performs. Lower is better because it means the design needs fewer resources overall. IEEE Spectrum reported that outsiders matched Google’s score within eight hours and beat it in about three days. By late July, more than 100 participants and their agents had brought it down to about 1.5 billion, roughly half Google’s figure. The leaderboard now reports a score 62.9% below Google’s.
I’d add a few caveats, most of them from Eigen’s write-up.
Nobody can use this to steal bitcoin. It would require a quantum computer far beyond those currently available, and the contest optimises only one part of the attack. Its test also supplies an input that a real attack would have to retrieve during the calculation. The researchers built another version that includes that extra work. It needed about 30% more of the expensive quantum operations, roughly 130 for every 100 used by the contest version.
Then there is where the idea came from. The same week the contest opened, a French researcher called André Schrottenloher published his own account of how Google had probably done it.
Eigen’s founder Sreeram Kannan told Spectrum he believes the agents read that paper and used it. So a person found the method, and the agents made it cheaper.
There was also no prize either. People probably joined because beating Google is fun.
If the technical bit lost you, this is how I would explain - A company puts a problem on Yukon, and people and AI agents try to solve it. Yukon checks their answers and shares each improvement so everyone else can build on it. The company gets more people trying to solve its problem without having to hire them all. People can use their AI agents to compete for rewards.
A team can build on someone else’s work and take the top spot, so Lighter counted each participant’s speed improvements across the whole contest. The six biggest contributors received ranked prizes, and three others won a draw weighted by their contributions. Earlier gains still counted, although the rewards measured improvement rather than how difficult the work was.
On 5 August, Lighter put the same loop on software it actually runs.
Lighter is a perpetuals exchange built as a zero-knowledge rollup. Every batch of trades has to come with a cryptographic proof that the exchange followed its own rules, and Ethereum checks that proof. Generating the proofs is slow. Lighter was doing it on about 3,000 Mac minis.
Lighter gave participants a challenge to make its software produce proof that 2,500 transactions had been handled correctly, but do it faster. The original version took about 13 minutes on a Mac mini. Every submission would run on the same type of machine and had to pass the same checks. Lighter initially offered 15,000 LIT tokens in prizes, worth roughly $35,400 that week.
Yukon tested each submission on one mac mini, then multiplied its speed by 3000 to estimate what Lighter’s full fleet could handle - those large numbers can be confusing, so a more useful comparison is how long one machine took to finish the job.
On the third day, an account called Meganpark980320 submitted a change written with OpenAI’s GPT-5 Codex. Yukon tested it on its own machine and confirmed that it still produced a valid proof. It completed the work about 19% faster than the previous best version. Participants could change only the permitted parts of the software, and Yukon calculated the score itself.
That improved version then became available to everyone. Within 75 minutes, two other accounts had made it faster again. One was using Anthropic’s Opus 5.
By the time the contest closed roughly three weeks after it began, the software was 10.57 times faster than the starting version. The job that originally took about 13 minutes now took roughly 74 seconds. Yukon listed 165 accepted improvements from 47 solvers.

I went through the public leaderboard to see how that improvement happened. Here are a few findings if you want the details:
Half the steps improved the previous record by about 0.6% or less. Only ten improved it by at least 5%.
Much of the progress came early. Within a little over four days, the software was already nine times faster. The remaining sixteen days brought it to 10.57 times faster.
The contributions were uneven too. Using a calculation that credits each account for its contribution to the overall speedup, five accounts accounted for roughly two-thirds of the improvement.
The largest share came from Gajesh2007, the account of Eigen Labs research engineer Gajesh Naik. This measures the gains attached to their submissions, rather than establishing who came up with every idea.
The record does show people building on one another’s work. In one 90-minute stretch on August 9, seven accounts using models from three AI companies set eleven successive records.

The improvements also reached Lighter’s actual service. Lighter says the first batch of changes reduced the time a machine typically took to produce a proof from about 133 seconds to 59 seconds. That is roughly 2.25 times faster in everyday use.
But the $35,400 prize was only part of the cost. Participants paid for AI, while Eigen ran the testing and Lighter’s engineers reviewed and integrated the changes. Those costs are not public, so we can see that the contest produced useful improvements without knowing how cheaply it did so.
Going back to Kaggle, and the Netflix Prize, which is the reason to be careful here. A merged team hit Netflix’s 10% target in 2009 and collected the $1 million. Netflix never used the winning entry. It wrote in 2012 that the extra accuracy “did not seem to justify the engineering effort needed to bring them into a production environment.”
Kaggle’s mechanics lean that way. A host posts data and a metric, and competitors are scored against test data they can’t see. Their work stays private unless they choose to share it, the contest ends on a deadline, and the top few take the money in exchange for licensing their entry to the host. Nothing in the rules makes today’s best entry tomorrow’s starting point. But Yukon’s rules do. The output is a change to the sponsor’s real code, already merged and already passing the sponsor’s checker, which is why Lighter could ship some of it eighteen days after the contest opened. That is a difference in mechanism, and Kaggle could copy it tomorrow.
Google bought Kaggle in 2017, and it already runs the same loop in private. Google DeepMind’s AlphaEvolve has a Gemini model propose code changes, scores them with an automatic evaluator, and keeps the best ones to mutate again. Google says it recovered 0.7% of its worldwide computing capacity that way. So the loop was already around when Yukon arrived. What Yukon adds is other people. Its bet is that outsiders, using different models and different ideas, can move a problem forward after one team’s agents have stopped, as ECDSA.fail did.
Crypto readers will be thinking of DeSci by now, so here is where it sits. Decentralised science changed who pays for research and who owns the result.
VitaDAO’s token holders vote on which longevity projects to fund and take a share of the intellectual property, and it says it has put $4.7 million into 31 projects, about seven average NIH grants. ResearchHub pays scientists $150 in its own token for a peer review. In both, the judging is still people reading a proposal or a paper, and VitaDAO’s money still goes out before the result exists.
Yukon has no token and no vote. It did not change who funds the work, instead it changed how the work is checked.
The names that come up most in AI-for-science are after a different kind of problem. FutureHouse, a nonprofit lab in San Francisco, builds agents that read papers and analyse data, and its commercial spinout Edison Scientific sells one called Kosmos. Edison says outside scientists checked its reports and found 79.4% of the statements correct. That still leaves about one in five wrong, and the customer has to find those.
Lila Sciences goes the other way and builds the checker. In drug or materials research, the only way to know if an answer is right is to run a physical experiment. So Lila has raised $550 million to build robotic labs where a model proposes experiments, machines run them, and the results feed back into the model.
Yukon only takes problems that are already easy to check, (meaning a machine can say yes or no, or give a score, without a person or a lab experiment) and pays a crowd to search for a better answer. Lila is spending hundreds of millions to build labs that can check an answer at all. They are not after the same problems yet.
My read on whether this travels beyond crypto and AI benchmarks is that it goes as far as cheap checking does.
Some of that is further than you’d expect. One closed Yukon contest, run with Carnegie Mellon, asked participants to make a common mathematical calculation less demanding. Computers often solve large sets of equations, including in software used to plan how electricity moves through a grid. Changing the order in which they tackle those equations can reduce the work involved. The winning entry needed about 16% fewer calculations than the standard method, according to Yukon’s scoring system.
Still, by my count, seven of the twelve contests on Yukon’s site are crypto-related. Crypto protocols have public code, a token to pay prizes in, and a zero-knowledge proof with its own checker.
Drug research is the hard case. The final check there is a clinical trial, and you don’t want to put that on the leaderboard.
A contest called CASP has scored protein-shape predictions since 1994, and DeepMind’s AlphaFold made its name there. If robotic labs like Lila’s ever get cheap enough to return a measured result for any molecule a stranger sends in, then in Yukon’s terms that lab is a checker.
Corporate R&D sits somewhere in between. Of the $937 billion the US spent on R&D in 2023, $625 billion was what the statisticians call experimental development, which means making known things work better.
Many engineering problems are harder to judge than wether software runs faster. A better battery, for example, might hold more energy but cost more to make or wear out sooner. Companies also have to decide what they are willing to share. Lighter’s improved code is public, so its competitors can inspect it too. A business that depends on keeping its methods secret may prefert to do that work intentionally.
For participants, there’s a catch. A researcher on a salary gets paid even when an experiment fails. AI only make it cheaper to try an idea. Someone entering a price contest will be spending time and money, sometimes to receive nothing. The sponsor gets to draw on that work without paying everyone who contributes.
The other shift is in who holds power, and it goes to whoever writes the test. Kannan told Spectrum he sees the scientist’s role as “architecting the right problem for a community of agents to make progress on.” For example, the ECDSA benchmark, participants were given a number that a real attacher would have to calculate along the way. The score left out the cost of that extra work.
Lighter checks whether the software produces proofs that Lighter’s existing checker accepts and how quickly it produces them. It does not also examine whether the rules being proved correctly cover everything the exchange should do. A submission can therefore pass the contest’s test without answering every question about the system’s safety.
Is this a market yet? In the Lighter contest the sponsor set a prize pool, and getting a contest listed starts with an application form. A market would let a company say what it will pay for each 1% of speed and let solvers decide if that’s enough. Nobody has built that part, as far as I can tell.
The nearest thing is Ridges, which runs on Bittensor. It is an open competition for AI coding agents, and “the highest-scoring agent earns emissions,” meaning a steady flow of newly issued tokens goes to whoever is on top. But nobody names what a given gain will pay, and Ridges scores general coding agents. A company can’t bring its own problem to it.
Yukon’s promise is that when a problem has a clear test, software can check the answer quickly and cheaply.
That leaves all the problems nobody knows how to score, which covers a great deal of science, and I do wonder what happens to that work once the money can go where the answer can be checked.
Token Dispatch is a daily crypto newsletter handpicked and crafted with love by human bots. If you want to reach out to 170,000+ subscriber community of the Token Dispatch, you can explore the partnership opportunities with us 🙌
📩 Fill out this form to submit your details and book a meeting with us directly.
Disclaimer: This newsletter contains analysis and opinions of the author. Content is for informational purposes only, not financial advice. Trading crypto involves substantial risk - your capital is at risk. Do your own research.










