I was reading through WRC 2026 in Beijing and saw lots of exhibitors on data collection for robots. Gloves, armbands, headbands, a whole aisle of motion capture (we're gonna call it mocap for the rest of the post) gear. I thought, huh, so mocap is the mainstream way of collecting robotics data?

Image from XinHua News
Why so much effort on collecting data for robots?
If you are trying to compare with LLMs, here's the difference. The Internet probably existed earlier than any of us remember. Nobody uploaded a photo or wrote a blog post with the intention of "this will be used to train an LLM in the future". It's just a byproduct of people living online for decades. And it just so happened that the language models got to eat them for free! (ok, not exactly free, but you get what I mean)
Robot data is a completely different case. It's not a byproduct of existing behavior; it requires the intention to create training data on purpose. And usable training data isn't a single video clip! It needs at least four things recorded at the same instant:
What it records | Why it's needed | |
|---|---|---|
Observation | What the robot sees right now | The input the policy reacts to |
State | What angle every joint is sitting at | Where the body is before it moves |
Action | Where every joint goes next | The thing being learned |
Force | How hard the contact was | Whether the glass survives |
If any of these four is off by just a few dozen milliseconds, the model will learn the wrong cause.
Every single data point also has to be a motion that the body can physically execute. You might have seen a lot of clips about this; every failure risks the hardware or the object. Someone has to demonstrate how a motion is exactly done.
Teleoperation
Most of the smooth demos you have seen are probably done by teleoperation.
Teleop means there's a human behind it, holding controllers or in a mocap rig, operating the robot in real time. This is the highest-fidelity data available in robotics, because the robots do the move and all the data points are recorded along with those motions. There's no translation layer in between.

Image from HTX
Teleop is also not as easy as it looks. The host from The Valley 101 (硅谷101) went to AgiBot's data collection factory in Shanghai and tried out the job of teleoperator. She failed a task that the trained collector completed in one pass. A skillful operator can have three times the output of a new hire. A new hire needs about a month to get good. You can tell this is a labor-intensive approach, despite the data quality.
This approach also has one critical ceiling. The operator can't feel what the robot feels. You are basically controlling another body from a distance, with no direct feedback.
From The Valley 101 program, the teleoperator also mentioned that an 8-hour shift for a professional can yield only two to three hours of usable data. They had to spend the other hours on scene setup, failed takes, etc.
Skild AI put it this way: "Mathematically unfeasible". Ain't nobody got time for that.
Can a robot learn from video?
Egocentric video (ego video) is a mainstream way to train a robot. It is video from a camera that you place on your head or chest. This is a first-person view because the framing & occlusions will all match what a robot sees.
And it scales. NVIDIA's EgoScale is pretrained on over 20,000 hours of first-person human video, spanning thousands of tasks and environments.
I learned this ratio across the industry from The Valley 101 show:
Teleop data to mocap data runs roughly 1:100. And mocap data to internet video data runs another 1:100. Generally, teleop data, which is the highest-fidelity one, is only about one part in ten thousand of the pool.
I mean, I knew it was the smaller slice. But I didn't expect four orders of magnitude. Two main things were missing for the remaining 9999/10000:
Force — It can see humans pick up the glass but can't see how hard they squeezed (or cracked it).
The body — A human hand has over 20 degrees of freedom. Most grippers have only one: open or closed. That thing you do with three fingers on the rim and a small rotation, the robot just doesn't have the body to do it.
This kind of mismatch is called the Embodiment Gap. One half is the visual gap, since the model only sees a human hand, but not the hand it will actually have. The other half is the state gap; mocap has to deal with self-occlusion and objects blocking the view. So whatever pose it recovers is already imprecise, which makes every action derived from it imprecise too.
Converting joint angles from the human to the robot is called retargeting. But robots are only imitating what humans do. They don't really understand what the motion was for. For example, if a human wants to pick up a pair of scissors, we have to push our fingers into the handles. If you only convert the angle of the finger joints and the positions of the fingertips for the robot, the motion will look the same, but the pressure applied to the handles will be very different, and the scissors would drop.
Universal Manipulation Interface (UMI)
What if the human holds the robot's hand instead of their own hands?
Good intuition. UMI is basically a handheld gripper on a stick with a camera mounted on it.

Image from Stanford Robotics Center
You hold it and walk around your own kitchen, do your stuff. The gripper's motion will then be recorded. This is open source, very cheap, and genuinely mainstream now.
But it only works for grippers. UMI records the motion of the tip. For a two-fingered gripper, that's pretty much enough. What if it's five-fingered like a human hand? The fingers also have joints, not just a tip.
DexUMI
DexUMI extends the UMI idea to dexterous hands.
You wear an exoskeleton that physically constrains your hand into the robot hand's kinematic structure. The robot hand can then make any motion that your hand can make. There's nothing left to retarget. For the video training, they would also replace human hands with robot hands (rendered in the video), so the model sees a robot working. Data collection is 3.2 times faster than teleop, with an 86% success rate across some tasks.
And… I think you can imagine the limitation. Every new robot hand needs its own exoskeleton, designed and tuned. Better, but still not very scalable right now.

Screenshot from BeingBeyond Youtube Channel
Back to the show floor at WRC
So what were those booths selling?
They are all chasing the same thing: to capture what a human does in a form a robot can use.
There are two I want to highlight that are going in opposite directions.
Tashan Technology
This is not a small player in the space. Their share in humanoid tactile sensors is around 80%. The approach is to put the sensor on the finger.
Weight is a major design constraint, and their TS-ECHO tactile fingertip is designed to weigh under 100g. Anything heavier might actually shift the hand's center of mass and eventually change how you move.
Tashan framed TS-ECHO as a tactile dimension on top of other vision-based collection approaches like NVIDIA's EgoScale.
SynapX
SynapX is looking at the same problem, but instead of hands, they have armbands that read EMG (electromyography) signals, which arrive before the motion. The force exertion is measured, and then the hand moves.
Which… I am genuinely skeptical about this. The muscle structure, fat thickness, and electrode placement differ for everyone. Although SynapX says its model now works on people it has never seen, with no new data and no fine-tuning.
Cross Embodiment
This is the term the industry uses for it: getting a skill learned on one body to work on another.
Teleop tries to skip this by having the robot do the moves itself.
UMI tries to do it on a smaller scale with only grippers.
DexUMI makes you wear the robot's kinematics.
Tashan tries to fill the gap that video can't carry: force.
SynapX tries to go upstream and read intent.
Every approach above is a different answer to it.
What the field can show today is probably object-level generalization. Like a robot that has seen similar objects can handle a new one, and environment-level generalization. Doing something it was never shown still has no credible evidence behind it.
Until robots are deployed at scale, there may be no way to find out.
The Hedge
Tashan opened a facility in Beijing's Shougang Park together with Professor Richard Sutton and called it a robot kindergarten in June this year. Robots accumulate experience through real physical interaction and reinforcement learning, explicitly modeled on AlphaGo Zero, which learned without human game records.

Image from 36Kr
From what I know, Professor Sutton has always been the guy with a strong conviction that imitation learning will be eaten by methods that learn directly from experience.
And Tashan sells fingertip sensors for capturing human demonstrations, then builds a lab with the man who thinks human demonstrations are a dead end.
Maybe nobody knows what the right answer is yet.
