Issue #214

JD.com's 150,000 Free Helmets — and One Hidden Camera

JD's free smart helmets for delivery riders hide a built-in camera, and I wonder who profits from that first-person footage.

SocietyJD.com's 150,000 Free Helmets — and One Hidden Camera

150,000 Free Helmets, and One Camera

On August 3, JD.com unveiled a smart helmet for delivery riders. It takes voice orders, analyzes veteran riders’ actual driving records to give directions, and sends a distress signal if the rider falls. The company says it will hand out 150,000 units free, on a priority basis, to its full-time riders.

Most coverage focused on “safety” and “efficiency.” But something else caught my eye: this helmet has a camera built in that reads the surrounding environment.

In the last issue, I told you about the textile workers in Karur, India, who strapped GoPros to their foreheads to film themselves folding clothes, earning around $2.6 an hour for it. A Chinese platform is now collecting footage of a similar nature — without paying anyone by the hour for it.

The smart helmets JD.com, Meituan, and Alibaba have rolled out over three years

Here’s what JD.com’s helmet can do, according to the company. Riders don’t need to tap a screen with wet hands — the helmet handles new orders by voice, lets them talk to customers hands-free, and reads out delivery notes like “please leave it at the door” automatically. Instead of relying on map apps, it learns the actual routes experienced riders have taken and gives turn-by-turn directions based on that. It even checks hygiene conditions automatically when riders pick up food at a restaurant. JD.com says it expects the helmet to cut navigation time by about 3 minutes per delivery, on average.

JD.com isn’t the one who started this. Meituan had already distributed more than 1 million smart helmets by June 2025. Theirs come with voice calling, fall-detection sensors, and automatic emergency response. Alibaba’s delivery arm (Ele.me at the time, now rebranded as Taobao Shangou) rolled out a similar helmet as far back as late 2022.

China had more than 10 million delivery riders as of early 2025, according to Xinhua News Agency. So over the span of three years, all three platforms have been pushing essentially the same kind of hardware onto this workforce.

The safety benefits these helmets provide are real. Handling things by voice is genuinely safer than fumbling with a wet smartphone screen in a downpour. Fall detection and one-touch emergency requests are features that can change survival odds when an accident happens. Strip away the marketing copy, and the effects are still there.

But the helmets from all three companies share one more thing in common: cameras, microphones, and location sensors fixed at head height, worn by riders crisscrossing the city all day long. And they’re all free.

The scarcest resource in robotics right now is first-person video

The scarcest resource in robotics right now isn’t models or chips — it’s video footage shot from the human eye level, in first person1.

Last June, researchers from Peking University, the National University of Singapore, MIT, UC Santa Barbara, and Nvidia published a paper called “HumanScale” that tackles this head-on. The experimental design is clean. They pretrained2 robots using the same 5,000 hours of data, split into two types: first-person video shot by humans wearing head-mounted cameras, versus driving and manipulation logs collected by actual robots.

The results ran against conventional wisdom. When tested on tasks never seen during training3, the model trained on human first-person video had roughly 20% lower loss. In physical robot experiments, the gap was even more dramatic. Robots with no pretraining had a 0% success rate on unseen tasks — robots pretrained on first-person human video hit 90%.

Why? The paper’s explanation is simple. Robot data is generated in labs, following fixed scripts, within the narrow reach of a stationary arm. Human video, by contrast, moves through homes, streets, and workplaces — the motion is fluid, and there’s little idle time. Even filming the same 100 hours, the number of motion trajectories extractable from human video was more than 5 times that of robot data.

So how scarce is this data, really? Ego4D, the academic standard for first-person datasets, contains 3,670 hours filmed by 923 people across 9 countries — a scale gathered over years through dedicated research projects.

Now put JD.com’s numbers next to that. Assume, very conservatively, that JD’s 150,000 full-time delivery riders each film just 1 hour a day. That’s 150,000 hours a day — the equivalent of 40 Ego4D datasets accumulating daily. That’s enough to fill HumanScale’s 5,000-hour pretraining set 30 times over, every single day. Scale that up to China’s full delivery workforce of 10,000,000 riders, and the numbers grow so large that comparing them to existing datasets stops being meaningful.

Let’s convert this using labor rates from Karur, a city in India. At $2.6 an hour, that’s $390,000 a day, and over $140,000,000 a year — for just 150,000 riders. That’s the amount of data-collection labor cost that never once showed up on anyone’s balance sheet.

For balance, though, two caveats are worth flagging.

First, the paper’s own limitations. As the authors themselves note, the comparison is capped at 5,000 hours (because public robot datasets of larger scale simply don’t exist). Physical robot validation was done on only 3 tasks and 1 robot platform. Whether the advantage holds at the scale of hundreds of thousands of hours remains unverified. That said, performance kept improving log-linearly all the way from 100 to 5,000 hours (R²=0.86~0.94), which reads as a signal that “more data keeps helping.”

Second, the nature of the data itself differs. What HumanScale dealt with was mostly hand manipulation tasks. What a helmet camera captures is closer to urban navigation — the layout of alleyways, shortcuts between apartment buildings, the route to find a store entrance, waiting in front of an elevator. This is raw material that maps more directly onto autonomous delivery vehicles and delivery robots than robotic arms. And, as it happens, autonomous delivery vehicles are exactly what’s spreading fastest across China right now.

The Line Item Missing from the Contract

The India case I covered last issue and this China case diverge in how the work gets paid for.

The model in Karur, India, had a name. There was a data-collection contract, an hourly wage, and workers who knew they were doing “AI training video filming” as a job. Even if the pay was unfairly low, the contract and the hourly rate were documented in writing. Because it was a transaction, it could be negotiated — and there was room for the rate to rise.

China’s model has no name. What the rider signed was a delivery contract; what’s on his head is safety gear. Data collection wasn’t a separate side gig — it was folded into work he was already doing. There’s no payment, no line item, and so there’s nothing to negotiate in the first place.

zingdungLast issue asked who owns AI training data. This time, I want to look at how you demand compensation when data collection isn’t written into the contract as paid labor at all.

The evidence, as it happens, is already sitting in the product manual. One of the flagship features of JD’s helmet is “veteran rider route guidance” — it analyzes the actual delivery records of top-performing riders and voices the optimal route to new riders. Look closely and this is a device that extracts a skilled worker’s experience into a model and redistributes it to unskilled workers. Nowhere does it mention paying the veteran separately for that know-how. This isn’t inference on my part — it’s the company’s own stated feature description.

The timing isn’t a coincidence either. On July 4, China’s State Administration for Market Regulation published a draft amendment to the E-Commerce Law. It would strengthen platform responsibility for the safety, wages, and basic social security of outsourced workers like delivery riders, and public comments were accepted until August 4. China has roughly 300 million flexible-employment workers.

But the riders’ reactions are telling. A delivery worker named Zhang Liang, quoted in an SCMP interview, works 12 hours a day and earns 300 yuan (about $44). What he said he wanted wasn’t social insurance — it was “higher per-order fees and reasonable delivery windows.” If insurance premiums go up, take-home pay goes down.

From the platform’s perspective, this legislation is a signal that labor costs are rising. Autonomous delivery vehicle costs, meanwhile, are falling fast. As of July, China had more than 15,000 autonomous delivery vehicles deployed nationwide, approved to operate in more than 200 cities. Vehicles that cost over 1 million yuan just a few years ago now sell for under 20,000 yuan. Hardware costs have dropped from $30,000 in 2021 to under $10,000 now.

JD.com Chairman Liu Qiangdong said at the APEC China CEO Forum in Beijing on June 21 that robots will take over deliveries, and that 700,000 delivery riders could ultimately be replaced. He added that the company is preparing job-transition training programs in partnership with more than 120 schools nationwide.

There’s a regulatory angle, too. In China, surveying roads and terrain and collecting geographic data isn’t something anyone can just do — it’s licensed, and geographic information collected by vehicles in particular is subject to separate controls. That means foreign companies are effectively locked out of scanning Chinese streets to build up their own datasets. So the city-scanning data generated every day above riders’ heads is, structurally, an asset that only domestic Chinese platforms can accumulate. The question I raised last issue — “who owns which layer” — reappears here, only at the level of the nation-state.

Let me be honest about where the line is. JD has stated it will use the data collected by the helmet to improve dispatch efficiency and safety management — it has never announced that the camera footage is used to train robots or autonomous vehicles. There’s no public evidence yet directly connecting the two. What I’m pointing to isn’t a causal claim — it’s the fact that the relevant pieces all sit inside a single company. Equipment that generates costly first-person video every day, a robotics business that needs exactly that kind of data, and a CEO’s own forecast that 700,000 delivery riders could be replaced — all of it lives inside JD.

Oswarld’s Lens

I think what matters more here than the surveillance controversy is which accounting line item data collection gets booked under.

There’s a pattern I’ve seen again and again while designing go-to-market strategies. When you want to cut a cost, instead of negotiating the unit price down, you simply stop listing that cost as a separate line item at all. In the Indian model, data collection was written in as a cost item. When it’s written in as a line item, someone can check that amount, dispute it, demand it be raised. In the Chinese model, the same activity gets folded into the safety-equipment budget. Nobody calls that expenditure a data purchase.

Labor issues become harder to fight not when the pay is low, but when the work itself was never treated as something exchanged for pay in the first place. If the unit price is low, you can point to that number and demand a raise. But if there’s no line item in the contract at all, you have no basis for saying what should be raised, or by how much.

And this isn’t a story about someone else’s delivery market. Domestically too, work apps, in-vehicle terminals, wearable scanners at logistics centers, and recorded customer-service calls generate massive amounts of work data every single day. Yet it’s rare for an employment contract or a subcontracting agreement to specify whether — and to what extent, and for what compensation — a company may repurpose that data to train AI.

When I help review technology adoption decisions these days, there’s one question I’ve started making sure to ask: “Does the contract include a clause on secondary use of the data this tool collects?” Most of the time, the answer is no. And “no” doesn’t mean it’s prohibited — it means nobody has ever asked the question.

I put out a weekly newsletter that cuts across technology, economics, and the humanities. Digging into questions like this is exactly what this newsletter is for.

Closing

This issue boils down to three things.

First, first-person footage shot at human eye level is the hardest training data to get in robotics right now. In the “HumanScale” experiment, pretraining on this kind of footage outperformed training on an equal amount of real robot data when it came to tasks the model had never seen before.

Second, in India, that footage was bought outright with hourly wages, while a Chinese platform is collecting the same kind of footage by embedding cameras in free safety gear. The difference between the two approaches isn’t really about the money — it’s about whether the data collection shows up as a line item in a contract.

Third, this helmet arrived at a moment when legislation strengthening riders’ social security protections was pushing labor costs up, just as autonomous delivery vehicle prices fell from the 1-million-yuan range to under 20,000 yuan. Both things are true at once: the helmet genuinely improves safety, and forecasts about riders being replaced have been stated publicly.

In the last issue, I asked whether the workers in Karur were being liberated or exploited. This issue’s question is a little different. What should we call this kind of data collection — one with no payment attached and no line item in any contract?

💬 Among the work tools you use right now — company apps, vehicle terminals, call recordings, wearable scanners — is there one that made you think, “the company could probably use this as training data”? Tell me in the comments which tool it was and why it struck you that way.


💬 Tell me in the comments about a tool that made you think “this could end up as training data.” I’ll factor it into the next issue. 📨 If you know someone who thinks about data, labor, and contracts, please pass this piece along.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • Ma, J., Bi, J., et al., “HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining”, arXiv:2606.20521, 2026. Link ··· This is the backbone of today’s piece. Start with the table comparing the diversity of human video versus robot data item by item, and the passage on real-robot success rates of 0% versus 90% — that’s where it clicks why egocentric video is so precious.
  • “HumanNet: Scaling Human-centric Video Learning to One Million Hours”, arXiv:2605.06747, 2026. Link ··· This is the original repository HumanScale drew its training data from. It gives you a sense of scale — over 800,000 of the 1,000,000 hours are egocentric footage.
  • “China’s tech giants race to put AI on delivery riders’ heads”, South China Morning Post, 2026. Link ··· An article comparing the helmets from JD.com, Meituan, and Alibaba side by side. This is where today’s piece started.
  • “Stronger social security? Proposed law sparks a pushback by gig workers”, South China Morning Post, 2026. Link ··· The key part is where riders themselves say they’d rather have per-delivery fees than social insurance. It shows the paradox of protective legislation.
  • “China’s Delivery Revolution”, The Wire China, October 2025. Link ··· Tracks in numbers how the price of autonomous delivery vehicles fell from ¥1,000,000 to ¥20,000. Read it alongside the labor-cost curve.

Background

  • Grauman, K., et al., “Ego4D: Around the World in 3,000 Hours of Egocentric Video”, CVPR, 2022. Link ··· The de facto standard dataset for egocentric video. Knowing the scale — 923 people, 3,670 hours — makes the arithmetic in this issue much clearer.
  • “‘700,000 delivery riders will no longer be needed’… China’s JD.com chief sounds the alarm”, Seoul Economic Daily, June 2026. Link ··· Lays out Richard Liu’s remarks at the APEC forum alongside the context of China’s 320 million gig workers.
  • Restrictions on geographic data in China, Wikipedia. Link ··· A quick overview of China’s surveying and mapping data regulations. Helps explain why this data becomes an asset exclusive to domestic platforms.

Past issues worth reading alongside this one


Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Egocentric video: Footage shot with a camera mounted at head or eye level. Unlike third-person video, it captures the exact moment hands touch objects and the actual angle a person sees from, which makes it far more useful for teaching robots how to move.

  2. Pretraining: A stage where a model builds up general competence from large volumes of generic data before being taught a specific task. It’s similar to learning how to hold a knife and handle fire before being taught how to cook.

  3. Out-of-distribution generalization: The ability to function properly in situations never encountered during training. This capability is what determines whether a robot can be useful in the real world, outside the lab.