Hello,
In 2017, a company called Kiwi started rolling out delivery robots on the UC Berkeley campus. You could order a burrito through the app, and a small bot would navigate the sidewalk to drop it at your door, all on its own. Students loved it, and Kiwi called it “parallel autonomy” because the whole setup genuinely felt like a glimpse of the future.
But turns out, every single one of those robots was actually being driven by a college student in Bogota, Colombia, using an Xbox controller for two dollars an hour. They had to feed the bot a new direction every five seconds, or it would just stop working.
That was 2017, and almost nothing has changed since, except that investors have now started pricing these companies like the next great SaaS. Let’s dig in!
Hermes Business: One Budget for Your Team’s AI Needs
Most teams pay for AI across different platforms. A few seats here, some tokens there, and a few API keys from multiple vendors that the company can no longer track. You may still be able to track the spending bill at the end of the month, but you will not know whether it’s paying off.
Hermes Business puts your whole team on one tab. Every member gets their own AI agent, reachable on Slack, Telegram, WhatsApp, or Discord. You can set how much each person can spend and track every dollar in one place.
Your team can pick from every model on Nous’s platform or bring their own keys.
The longer your team uses it, the more it knows your team members.
Need the agent to run on your own servers? Hermes Enterprise runs on infrastructure you control.
Offshoring Autonomy
Japan has the most severe labour shortage in the developed world. A third of its population is over sixty-five, its working-age workforce peaked in 2017, and the country now faces an estimated shortfall of eleven million workers by 2040.
We see that by 2040 there will be 3 humanoid robot startups per man!
Convenience stores across the country simply cannot find enough people willing to stock shelves. So if there is a single place on the planet where fully autonomous robots should work, where every demographic and economic pressure is aligned to make them work, it is Japan.
And they are actually moving towards it quite fast, ironically. Engineering graduates in a Manila office building wear VR headsets and remotely restock shelves in over 300 FamilyMart and Lawson stores across Tokyo. Telexistence, a Tokyo startup, built the robots. The operators come from Astro Robotics, a Filipino workforce contractor that pays them between $250 and $315 a month, and each worker monitors around fifty robots at a time, waiting for one to drop a can or misjudge friction on a wet bottle.
When that happens, they strap on the headset, take manual control of the gripper, and nudge the product into position. This happens about 50 times during their 8-hour shift, and it has even caused workers dizziness and blurred vision from constantly switching in and out of VR.
Thankfully, the robots can handle about 96% of tasks on their own. The AI does the scanning, the routing, the basic pick-and-place. But the remaining 4% is why a human being sits in a chair in Manila. That 4 per cent might sound small, but it is where the industry’s economics live.
Think about what actually goes wrong during those fifty interventions per shift. A robot picks up a dry bottle from a standardised shelf a thousand times without issue, because the AI was trained on exactly that scenario. Then it hits a bottle with condensation on it, and the grip slips, because wet plastic has different friction than dry plastic, and the model can’t account for that. Or a dented can shows up that does not sit in the gripper the way the training data predicted, or a new product arrives that the system has never seen, so there is no template for how to hold it. There’s literally a long tail of weird shit that the real world might throw at a robot!
And each of these situations requires the robot to read an unfamiliar physical situation and figure out, on the spot, how to respond with its hands. A human does this unconsciously, but a robot has no way to. When the robot encounters one of these cases, it stops automatically and flags the failure to a human in Manila, who takes over remotely through a VR headset, using the robot’s cameras as eyes and its gripper as hands. The operator then fixes the problem, releases control, and the robot continues on its own until the next failure.
This setup, where a remote human steps in to cover the AI’s blind spots from thousands of kilometres away, is what the industry calls teleoperation. It is the operating layer that keeps the entire robotics business running right now, and every company in the space depends on it to some degree, whether they say so publicly or not.
The alternative is to make the robot smart enough that it never needs to call Manila. For that, the robot would need to do three things at once.
See what is in front of it through its cameras
Figure out what to do about it (squeeze harder, approach from a different angle, try again)
Physically execute the movement, all in a continuous loop fast enough to keep up with the real world.
You are essentially trying to give a machine the kind of hand-eye coordination that a two-year-old already has, and it turns out that requires running very large AI models directly on the robot’s own hardware. The industry calls these vision-language-action models, or VLAs. A VLA processes the camera feed, interprets the scene (that bottle is wet, that shelf is crooked, that product is unfamiliar), decides on a physical response, and sends commands to the gripper, all in milliseconds. If the model takes even half a second to figure out how hard to grip a wet bottle, the bottle is already on the floor.
Running a VLA on a robot means putting a chip inside the machine that can handle all that computation locally, without sending data to a cloud server and waiting for a response.
And that’s why a robot in a Tokyo convenience store cannot stream video to a server farm in Virginia every time it needs to decide how to grab a can. The processing has to happen on the robot itself, in real time. The dominant chip for this right now is Nvidia’s Jetson AGX Thor; it is currently the only commercially available processor that can run VLA models at the speed and performance that real-world physical manipulation demands. When your robot needs sub-100-millisecond decisions about grip pressure and approach angle, the hardware options shrink to essentially one company. Every major robotics startup building toward autonomy is building on Nvidia’s silicon, which gives Nvidia the pricing power that comes with being the only supplier in a market where everyone needs what you sell.
And Nvidia charges accordingly. The Jetson Thor developer kit costs $5,499 as of August 2026, up 57% from the $3,499 Nvidia quoted when it first announced the chip. The price keeps climbing because VLA models keep getting bigger with each generation (better manipulation accuracy requires more parameters, which requires more compute to run) and because nobody else makes a competitive alternative.
Amortise that $5,499 over a three-year useful life at reasonable utilisation, and you are paying somewhere between 20 and 40 cents per robot-hour, just for the hardware, before you account for the software stack, the systems integration, or the engineering team keeping the models stable.
Now let’s compare that to what the human costs are:
What’s most interesting is that the cost gap between these two columns is growing in the wrong direction. The human side keeps getting cheaper because operators keep watching more robots at once. Kiwi needed one person per robot in 2017. Astro has one person per fifty in 2026.
That ratio will keep improving as monitoring tools get better and latency drops. At the same time, the silicon side keeps getting more expensive because VLA models demand more compute with each generation, and Nvidia has no competition pushing prices down. Five years ago, you could at least argue that compute costs would eventually fall below human labour costs. Today, that crossover point is further away than it has ever been.
When a company’s cheapest and most reliable input costs 3 cents per robot-hour, and the technology that would replace it costs ten times more and still cannot handle the hard cases, every rational dollar goes into making the cheap input even cheaper.
Now you can use the saved margins to build better monitoring dashboards, lower-latency headset connections, and faster pilot onboarding programs. No company would ever spend that money finishing the autonomy stack. A company called Adamo was recently launched with this exact logic. Adamo calls itself “the bridge to robot autonomy.” It sells managed teleoperation services to US robotics companies at $13 an hour.
Adamo’s own sales materials make the case that building an in-house teleoperation team costs more than outsourcing to Adamo, which means the company is actively selling the proposition that the human should stay in the loop and that you should pay Adamo to manage them. And every efficiency gain Adamo makes in pilot training or shift scheduling makes that spread more profitable and makes full autonomy less economically necessary.
And not just that: Telexistence, the company running robots in those 300-plus FamilyMart stores, adds another darker layer. Every time a Manila pilot takes manual control and corrects a robot’s grip, that movement gets logged as embodied teleoperation data. The company then feeds this directly to Physical Intelligence, a San Francisco AI lab building VLA models that are supposed to eventually make teleoperation unnecessary.
So the workers are doing three things at once: performing the physical labour that keeps the stores stocked, keeping unit economics attractive enough that investors keep funding the operation, and generating the training data that will theoretically make their own jobs obsolete.
These workers are building the tools that will eventually replace them. And because the economics are so favourable, the workers keep working, caught in a loop that benefits every participant in the chain except the person wearing the headset.
In February 2026, Senator Markey’s office wrote to seven autonomous vehicle companies, including Aurora, Tesla, Waymo, and Zoox, asking how often their remote operators intervene. Every company refused to answer. Before GM shut it down, Cruise had one remote assistant for roughly every fifteen to twenty vehicles, with human help triggered about every four to five miles.
The $12 Million Robot
So the teleoperator costs 3 cents an hour, and the chip that would replace them costs ten times more and still cannot do the job properly. And the companies making these robots have every incentive to keep the human in the loop. Given all this, you would expect the market to price this accordingly? Well, good luck!
Figure AI, the humanoid robotics startup, raised at a $39 billion valuation earlier this year. The company has shipped roughly 150 robots. And you would be surprised that the top five Western humanoid startups have a combined valuation of $73 billion. Together, they have deployed fewer than 400 robots. Which, if you do the math, works out to about $12 million in invested capital per robot that actually exists in the physical world.
For Figure’s valuation to make sense at even a 15x revenue multiple, the company would need to reach around $2.6 billion in annual recurring revenue. At a thousand dollars a month per robot, which is the standard Robot-as-a-Service price point the industry has converged on, that would REQUIRE ROUGHLY 215,000 robots deployed.
And even if you could manufacture and deploy 215,000 humanoids tomorrow, the unit economics still wouldn’t work. A humanoid robot costs at least $60,000 to build. Over a three-year useful life, that is $20,000 a year in hardware depreciation alone. Add $4,000 a year for maintenance and repairs (service every six months, shipping, technician time, lost revenue during downtime) and another $2,000 for remote supervision, and you land at $26,000 per robot per year in total operating costs.
On the value side, a robot replacing a $30-an-hour warehouse worker across six shifts a week produces about $75,000 in annual labour savings. That sounds like a great margin until you account for two more things. Robots currently work two to four times slower than humans at the same tasks. And real jobs are not single tasks; they are dozens of tasks, and a robot trained on picking will sit idle when the work shifts to inventory or packing.
Apply a conservative 3x productivity penalty and a 15% flexibility discount, and the $75,000 in value drops to about $21,000. Subtract the $26,000 in costs, and the robot is negative $5,000 a year.
Negative 18% ROI.
This is a technology that is genuinely impressive in demos, with a TAM slide in pitch decks that assumes total addressable penetration, for a valuation that prices in a future where every technical and economic obstacle has been solved simultaneously, and unit economics that are underwater today.
The narrative runs ahead of the fundamentals; the fundamentals take longer than anyone projected, and the correction, when it comes, is priced in all at once.
That’s all for today!
Vaidik
Token Dispatch is a daily crypto newsletter handpicked and crafted with love by human bots. If you want to reach out to 165,000 subscriber community of the Token Dispatch, you can explore the partnership opportunities with us 🙌
📩 Fill out this form to submit your details and book a meeting with us directly.
Disclaimer: This newsletter contains analysis and opinions of the author. Content is for informational purposes only, not financial advice. Trading crypto involves substantial risk - your capital is at risk. Do your own research.











