Chapter 7: Human Demonstrations — Translating Work into Robot Action
Overview
Human work is abundant, but it does not automatically become robot action. Human hands have more degrees of freedom than many robot hands, regulate force through skin and muscle, and hide expert intent and near-damage states inside motion that video alone cannot explain. The chapter’s question is therefore not simply whether teleoperation can be eliminated. It is which transfer losses and validation debts appear when the target robot is released from collection duty.
Direct teleoperation provides commands that were actually executed on a robot, but occupies the robot and requires resets. Handheld, wearable, and egocentric collection can spread across worksites before a robot cell exists, yet must reconstruct coordinate frames, viewpoint, timing, contact, and quality outcomes. UMI made this trade explicit with a handheld two-finger gripper, relative trajectories, and latency matching [1]. FastUMI and YUBI scale interface construction and collection; DexCap, DexUMI, and DEXOP capture more of the hand; EgoMimic and EgoDex exploit egocentric observation [30], [31], [29], [3], [4], [8], [36]. These are not interchangeable replacements for teleoperation. They are measurement systems that choose different losses.
After reading this chapter, you will be able to... - Compare direct teleoperation and robot-free collection using robot occupancy, operator attention, reset labor, calibration, and accepted-episode yield. - Diagnose five transfer losses: kinematic feasibility/workspace, viewpoint/occlusion, action/timing, force/touch/contact, and task outcome/quality. - Explain which losses UMI, FastUMI, YUBI, DexCap, DexUMI, DEXOP, and EgoMimic reduce and which debts they retain. - Distinguish inspectable paper evidence from first-party company claims about gloves and capture platforms. - Design real-robot replay, quality labels, worker rights, and update lineage before human-data scale becomes manufacturing evidence.
Five Transfer Losses and the Denominator Problem
Human-to-robot transfer must evaluate kinematic feasibility and workspace, viewpoint and occlusion, action space and timing, force/touch/contact state, and task outcome and quality as separate losses. A low loss on one axis does not compensate for a product-damaging loss on another. This taxonomy is bounded to the tasks, embodiments, sensors, controllers, data splits, and evaluation denominators in the cited work; it is not a universal ranking [32], [3], [33].
| Transfer loss | Direct teleoperation | Handheld, wearable, or egocentric capture | Release question |
|---|---|---|---|
| Kinematic feasibility and workspace | Commands are executed on the target morphology, but leader hardware and limits are arm-specific [27]. | UMI/YUBI share grippers, DexUMI constrains the hand, and DexCap solves inverse kinematics. | Do joint margins, self-collision, fingertip material, and arm dynamics pass in the real cell? |
| Viewpoint and occlusion | Robot cameras resemble deployment views, although the robot and hand still occlude contact. | Egocentric video scales broadly, but human-hand appearance, camera height, and backgrounds differ. | Does hand inpainting or view alignment survive reflective, transparent, and fixture-confined scenes? |
| Action and timing | Joint/end-effector commands and timestamps are present, while display and communication delay enter operator behavior. | Video and hand pose require action inference; UMI must match collection and deployment latency [1]. | Are policy rate, action chunks, stops, and safety-controller semantics equivalent? |
| Force, touch, and contact | Robot force can be logged when instrumented, but haptic feedback and sensor placement can differ. | RGB/pose data contains no calibrated force; UMI-FT, OSMO, and RealDexUMI add partial contact channels [2], [38], [32]. | Do units, bandwidth, mounting, wear, and recalibration retain meaning across capture and rollout? |
| Task outcome and quality | Robot execution links readily to success, but binary success hides damage, rework, and cycle-time tails. | Human video can show intent and outcome, not robot executability or inspection acceptance. | Are accepted yield, defect codes, rework, protective stops, and delayed inspection joined to the episode? |
The table breaks apart the claim that more human data is enough. Some data preserves hand geometry, some preserves force, and some mostly broadens observation. The axes should not be collapsed into one score: a visually plausible trajectory may exceed joint limits, and a kinematically accurate insertion may still apply damaging force.
Figure 7.1. Author-created comparison of how UMI, DexUMI, DEXOP, and EgoScale move human work toward robot-executable data through retargeting and release checks. It is not a shared experiment that quantitatively compares transfer loss across the pathways.
Historical Baseline: Direct Demonstration Occupies the Robot
The historical baseline for robot-learning demonstrations includes kinesthetic teaching, leader–follower devices, and VR teleoperation. MIME paired human videos with robot trajectories, while RoboTurk used commodity mobile sensing and cloud teleoperation to crowdsource 6-DoF trajectories [25], [26]. GELLO built a low-cost leader with the target arm’s kinematic structure, and ALOHA combined bimanual demonstration hardware with ACT’s chunked action prediction [27], [28]. The advantage is concrete: joint state, end-effector pose, gripper command, and robot-camera observation share a timeline, so the identity of the executed action is relatively unambiguous.
One hour of direct demonstration, however, is not merely one operator hour. A target robot or matched collection robot is occupied. Someone must reset parts and fixtures after failure, monitor collisions, and screen incomplete episodes. A leader built around a particular arm incurs construction and calibration again when the arm changes. Successful-demonstration counts in RoboTurk, per-task minutes in ACT, and trajectory counts in fleet datasets use incompatible units. A total cost denominator must include robot occupancy, active operator control, passive monitoring, reset labor, rejected episodes, and later validation [24].
Direct robot-native collection remains a strong baseline for contact-rich, narrow-tolerance work. Joint limits and safety controllers operate while the robot approaches the fixture, and an instrumented robot can record force in its own frame. Yet teleoperation without haptic feedback can delay recognition of slip and jamming, and the operator’s compensation for display or network latency becomes part of the data. Robot-native data is not perfect truth; it is simply the clearest evidence of which command was physically executed.
This baseline makes the benefit and debt of robot-free collection legible. A portable interface can visit many workstations while the expensive robot remains available for production or validation. The new transfer stage then consumes tracking, retargeting, calibration, and robot replay. The proper efficiency metric is not captured episodes / operator hour; it is real-robot accepted episodes / total operator, robot, reset, calibration, and validation time. Few papers disclose that full denominator.
UMI Lowers the Cost of Accessing the Worksite
UMI-style interfaces remove the target robot from collection but remain dependent at deployment on tracking quality, matched capture/deployment latency, and actions feasible for the robot [30], [32], [3]. UMI asks a person to manipulate a handheld, camera-equipped two-finger gripper. Relative trajectories reduce dependence on a global frame, and latency matching addresses the observation–action timing shift [1]. For manufacturing, the important benefit is process discovery before a robot cell exists: task order, approach directions, and layout variation can be observed without allocating a robot fleet.
FastUMI decouples specialized UMI components and uses off-the-shelf tracking, reporting more than 10,000 trajectories across 22 everyday tasks [30]. The number establishes dataset scale and scope, not the fraction of trajectories accepted in real-robot replay or the total reset and calibration time. “Hardware-independent” also does not mean controller-independent. Velocity, acceleration, safety limits, and control rates still need target-specific integration.
YUBI extends the shared-interface principle with finger-aligned yielding grippers, VR 6-DoF tracking, and the same gripper mounted on multiple arm families [31]. Its technical report lists 8,434 hours, 1.20 million episodes, 119 tasks, 179 operators, and 22 collection stations. Those are strong first-party research measurements of collection scale. They are not accepted-production yield, rejected-episode rate, long-term drift, or independent factory validation. Sharing the gripper avoids part of the embodiment problem rather than solving transfer between arbitrary hands.
UMI does not remove every loss. A two-finger gripper interface fits many manufacturing tasks, but flexible cable routing, cloth-like material handling, and finger-gait manipulation may lose too much structure. UMI-FT matters because it adds force/torque and therefore preserves part of compliant behavior [2]. Without force, “the insertion succeeded smoothly” and “the worker forced it through” can look similar on video. Its evidence is nevertheless bounded to three single-arm tasks and the reported sensor configuration; tethering and sensor-delamination risk also belong in the operating cost.
DexCap Frees the Hand and Adds Retargeting Debt
DexCap combines portable SLAM with electromagnetic hand tracking, which is less vulnerable than pure vision to finger occlusion [29]. DexIL then uses inverse kinematics and point-cloud imitation to map human hand/object motion to a robot hand, with optional human correction during rollout. The architecture shows that robot-free collection and selective correction are complements: broad human motion can initialize the policy, and robot-side correction can target states where transfer failed.
Electromagnetic tracking does not measure contact force. Inverse kinematics can match fingertip locations while losing joint margin, self-collision feasibility, friction, or palm support. Selective correction also appears cheap if a paper counts only corrected segments. Its true cost includes the time an operator watches the policy, takeover latency, safety stops, and resets. Each correction should record why intervention occurred, whether it was residual or full takeover, pre-intervention risk, and subsequent outcome [40].
DexUMI and Retargeting Preserve Motion While Changing the Hand
DexUMI treats the human hand as a universal manipulation interface and transfers motion to a dexterous robot hand [3]. A wearable exoskeleton constrains human motion toward robot-feasible kinematics, while robot-hand inpainting reduces the visual gap. That matters because it tries to preserve in-hand adjustment, multi-finger contact, and regrasping. The reported 86% average success across its two-platform suite belongs to those tasks and denominators; it is not a universal transfer rate for arbitrary hands and camera layouts.
Figure 7.2. DexUMI Figure 1 shows exoskeleton demonstrations, robot executions, and long-horizon, contact-rich, multi-finger, and precise task examples on XHand and Inspire Hand. The teaser itself does not establish preservation of contact force across different hand structures. Source: Xu et al. 2025, arXiv:2505.21864 Fig. 1.
Retargeting is interpretation. Human joint ranges, skin friction, fingernails, fingerprints, and felt pressure differ from robot hands. A retargeted trajectory can be kinematically plausible while failing at force closure or tactile state. The exoskeleton improves feasibility by changing natural human motion, and per-hand fabrication and calibration remain costs. Robot-hand inpainting also requires real hardware data and has not established camera-layout generality. DexUMI-style data must therefore be evaluated with robot-side execution logs, contact failures, and calibration state, not just hand pose.
RealDexUMI uses the same lightweight dexterous hand, in-hand camera, and fingertip tactile module during wearable collection and deployment [32]. Its reported 88.75% average across eight tasks supports the promise of shared end-effector design. “Zero gap,” however, can only describe the shared end effector. Arm dynamics, external viewpoint, controller delay, fixture compliance, and task-quality labels remain different. A small preprint task suite does not establish long-shift manufacturing reliability.
Decision Walkthrough: Adhesive Gasket Placement
Consider placing an adhesive gasket along the edge of a metal housing. A worker holds the gasket, peels part of the protective film, anchors one corner, and controls tension while pushing out bubbles and wrinkles. On video the hand may look as if it simply traces a smooth loop. The real skill is more detailed: when the adhesive can no longer be lifted, how much pressure avoids cosmetic defects, and how to recover when the liner sticks to the finger.
UMI-style two-finger data can quickly capture approach poses and the coarse path [1]. But the contact sequence that presses the gasket edge into place is easily lost in two-finger action. DexUMI-style retargeting preserves more finger sequencing, but pressure changes when the robot fingertip compliance differs from human skin [3]. DEXOP-style devices can observe contact more directly, but only if workers can wear them for long periods [4].
The walkthrough shows that human-to-robot transfer is three validations, not one conversion. First, did human capture observe the relevant skill? Second, did retargeting produce the same contact intent on the robot morphology? Third, did robot rollout satisfy inspection and downstream quality? Knowing which stage failed determines whether to collect more data, change the sensor, or change the hand.
Retargeting Validation Matrix
Visual retargeting and inverse-dynamics actions must be validated by closed-loop real execution rather than appearance similarity alone. The requirement applies to RoboPaint’s scene-reconstruction route, Human2Sim2Robot’s human-to-simulation-to-robot route, and HumanPlus-style human-motion transfer, while remaining bounded to each paper’s tasks, cameras, embodiments, and evaluation denominators [34], [33], [35]. Plausible-looking hand/object motion can hide reflective or transparent reconstruction errors, occluded contact, bad depth, and action–state inconsistency.
| Validation layer | What to check | Where a failure sends the team |
|---|---|---|
| Kinematics | Fingertip position, joint margin, self-collision, robot-hand revision | Retargeting rule or hand structure |
| Contact and timing | Contact order and duration, slip, pressure proxy, policy rate, action chunks, stop/restart | Fingertip material, tactile sensor, controller mode |
| Outcome and operations | Quality label, cycle-time tail, skilled intent and recovery | Data selection, inspection rule, task allocation |
DexUMI widens the bridge between human hand pose and robot-hand action, but also creates more audit points [3]. A human thumb-index pinch can produce a different force distribution on a robot hand, and palm support can become slip on a robot palm. Converted data should become a production candidate only after all three layers pass. The matrix also separates tasks that need a five-finger hand from those a two-finger gripper can handle: the goal is not a human-looking robot, but evidence for the hand structure each task needs.
Human-Video Scale Without Force Labels Creates Overconfidence
The scale advantage of EgoMimic and EgoScale is real [8], [9]. Seeing many workers handle many SKUs expands the visual prior. The system can learn where tools are placed, how hands approach, and where workers pause their gaze. But video does not know force. It does not reliably tell whether the worker pressed lightly, forced the part, damaged a cable jacket, or caught slip through the fingertips.
UMI-FT and CoinFT attach force channels to this scale problem [2], [12]. What matters is not simply that a force sensor exists. The force trace must have the same meaning across human capture and robot rollout. If the human device and robot end-effector have different force ranges, thresholds shift. A force label is interpreted through calibration, mounting, and contact surface, not as a raw number alone.
Manufacturers should use human-video pretraining as a wide prior and force-aware robot rollout as release evidence. Good video scale should not lower the production gate. Conversely, force-only data that is too narrow will miss SKU variation. The two data types are complementary: human video proposes behavior, while force/tactile robot data verifies that the behavior does not damage the product.
How to Read Tactile-Glove Evidence
Tactile gloves invite overinterpretation because “giving the robot what the human felt” sounds direct. STAG used distributed palmar pressure cells to study human grasp signatures, but did not cover every hand surface, its piezoresistive elements had hysteresis and drift, and the signals were relative pressure rather than calibrated force [37]. The work provides an inspectable paper, device description, and experiment. Its object-classification results still cannot be converted into robot-policy contact success.
OSMO Tactile Glove is a paper-based attempt to align human and robot tactile hardware more directly, while UniTacHand explores cross-hand tactile representation [38], [39]. Their evidence is bounded to a human/robot-hand design or a small paired dataset and one platform. Retargeting and calibration must be redesigned for dissimilar end effectors. With those boundaries, a glove can be evaluated as an instrument that reduces contact loss, not advertised as a universal behavior converter.
Company product pages or promotional videos establish the vendor’s disclosed configuration and first-party claims—such as “high resolution,” “natural haptics,” or “instant transfer”—but not independent task performance. Even sensor count, sample rate, and wearable duration are not comparable without protocols. Procurement should request raw-data access, unit-bearing calibration curves, drift after repeated donning, cleaning/replacement reproducibility, and real-robot evaluation. Company glove claims and task success from inspectable papers should never occupy the same performance leaderboard without a shared test.
Operating Model for Worker Data
A human-data system needs an operating model. Who wears the device? Which shifts collect data? How are expert and novice workers separated? What happens when a worker asks for deletion? Without those rules, data quality changes with shift, supervisor, and vendor operator.
Consent is not a checkbox; it is a feedback loop. Workers need to know why data is collected, what is anonymized, and how model updates may change their work. Process-IP protection needs the same care. Background fixtures, jigs, custom tools, and inspection sequences may reveal competitive knowledge. Excessive redaction removes learning context; weak redaction leaks IP.
Vendor rights must also be explicit. Even if a vendor processes raw episodes, the manufacturer needs exportable lineage. After a model update, the team should be able to trace which human segments contributed to improvement. Otherwise the human-data flywheel becomes the vendor's black box rather than the manufacturer's asset.
DEXOP and Exoskeleton Demos Observe Contact More Directly
Exoskeleton and direct-contact devices can reduce kinematic mismatch while sacrificing universality, comfort, or per-hand portability [5], [4], [35]. DEXOP uses a passive exoskeleton to capture direct-contact dexterous demonstrations. It records vision and tactile data and provides direct contact and force feedback, but does not establish perfectly calibrated force preservation across arbitrary hands. ExoStart’s simulation amplification from a small exoskeleton seed belongs to the same line, while its simulation/real control-rate mismatch, soft-contact failure, and vision-student gap remain counterevidence.
Figure 7.3. DEXOP Figure 1 shows the wearable device with whole-hand tactile sensors, precise perioperation, and bimanual, long-horizon, contact-rich manipulation examples. The figure does not quantify wearability, repeatability, or worker burden. Source: Fang et al. 2025, arXiv:2509.04441 Fig. 1.
In a factory cell, technical feasibility is not enough. The team must check whether workers can wear the device for long periods, whether the glove satisfies safety rules, whether cleaning and ESD requirements pass, and whether data collection disrupts cycle time. Exoskeleton demos are expensive labels, so they should target contact-critical segments rather than every operation.
Tier Data Collection to Match Cost and Risk
Human data collection should not depend on one device. Matching cost and risk across the following tiers makes PoC design and procurement criteria explicit.
| Collection tier | Primary use | Evidence to add before release |
|---|---|---|
| Tier 1: video or UMI-style capture | Broad task variation and visual routines [1] | Force/tactile labels for contact-critical segments |
| Tier 2: UMI-FT, CoinFT, or tactile patches | Partial observation of force-sensitive segments and slip [2], [12] | Sensor calibration and real-robot force safety |
| Tier 3: DEXOP or exoskeleton capture | Precise capture where finger-contact order is the skill [4] | Wearability, process safety, and robot-side quality replay |
Capturing every operation at Tier 3 is slow and expensive; capturing everything at Tier 1 misses contact failure. Sensitivity also differs: egocentric video has a broad privacy surface, exoskeletons record worker motion deeply, and force/tactile sensing may reveal process know-how. Retention, access, and redaction should therefore differ by tier, while every tier converges on the same robot-side QA replay.
Choosing not to collect is also a design decision. Rare repair, sensitive process IP, or a conflict between wearable sensing and safety rules may favor a controlled robot-side experiment. A large-data strategy identifies which data can change a policy update; it is not a strategy of recording everything.
This avoids a common governance failure. Teams sometimes collect broad human video because it is easy to justify, while the actual blocker is a narrow contact segment with no force label. The result is a larger archive and the same deployment risk. Data collection should begin from the failure that blocks release, not from the camera that is easiest to mount.
The collection plan should therefore name the expected policy change. If a dataset is supposed to improve pregrasp pose, it needs different labels than a dataset meant to improve slip recovery. If no one can say which model decision the new human data will change, the collection should pause. This discipline prevents human-data programs from becoming storage projects.
It also helps procurement. A vendor offering egocentric-video scale, a glove system, and a force-sensing interface is not selling three versions of the same data. Each product occupies a different tier in the failure map. The manufacturer should buy the tier that matches the current release blocker, then require robot-side replay evidence before expanding collection.
The same tiering should appear in budgets. Cheap video can be collected broadly, but expensive contact capture should be reserved for failures that materially block release. Otherwise the data program spends the most money where the model gets the least new signal.
Egocentric Video Gives Scale, Not Force
Egocentric human video can scale observation collection, but it does not directly provide robot actions, forces, or guaranteed executability. This boundary applies to EgoDex’s passive video scale, EgoMimic’s aligned human/robot co-training, and HumanPlus-style shadowing, and must remain bounded to their tabletop bias, learned alignment, or evidence from a customized 33-DoF humanoid and its reported task suite [36], [8], [35]. Skilled workers’ gaze, hand approach order, and tool-use patterns can be collected without robot occupancy. That is valuable in high-mix manufacturing, where many SKU variations must be seen before the cell is finalized.
Egocentric video does not directly provide contact force or quality outcome. That is why vision-touch representation work such as [10] and [11] matters. Human-video pretraining can create a visual prior, but insertion jams, over-force, and slip recovery remain weak without force/tactile data or inspection labels. Egocentric data should therefore be treated as cheap pretraining, while release evidence comes from robot-side replay and QA traces.
Worker Consent and Process IP Are Part of the Pipeline
The moment human data is collected, manufacturing automation brings privacy and IP into the learning pipeline. Worker video is personal data, hand motion is skilled know-how, and the background fixture or tool layout may be process IP. Consent, redaction, retention, role-based access, and vendor export rights must be decided during collection design, not after model selection.
This is a quality issue as much as an ethics issue. If privacy rules delete important override segments, edge-case labels disappear. If process-IP rules prevent a vendor from seeing any raw episode, model improvement may slow. Conversely, if the vendor controls every raw episode and update lineage, the manufacturer loses the reason and reproducibility of improvement. Governance of the human-data pipeline directly affects performance.
Manufacturing Cell Checkpoint
First, divide the task by transfer loss. Pick-and-place tasks that fit two-finger grippers can start with UMI-like data. Tasks where finger contact order is central should consider DexUMI or DEXOP-like capture. Tasks where force decides quality should include UMI-FT, force/torque, or tactile patches from the beginning. Tasks needing large visual variation can use egocentric video pretraining as a candidate.
Second, evaluate human data and robot execution separately. Human-side success is not robot-side success. After conversion, robot rollouts must be checked on the same SKU, lot, fixture split, and quality labels. Third, encode worker consent and process-IP protection into the data schema. The collection spec should state who can view raw data, what is anonymized, and whether the manufacturer can export raw episode lineage after model updates.
Counterevidence, Limitations, and Missing Denominators
First, EgoMimic reports favorable marginal value for added human hours in its aligned co-training regime, but that is not a universal exchange rate between human and robot hours. Strong UMI, DexUMI, and YUBI results also show that transfer often depends on engineered constraints: a shared end effector, an exoskeleton, hand inpainting, or latency matching. Abundant human video and executable supervision are not interchangeable.
Second, large collection claims often omit accepted-episode yield. YUBI’s reported hours, episodes, tasks, operators, and stations establish collection scale, not factory defect rate, reset labor, maintenance, or independent reproduction. DexUMI and RealDexUMI averages require remeasurement when task, platform, camera, or end effector changes. The public evidence does not establish that robot-free capture universally lowers total data cost.
Third, retargeting can create kinematic success without contact success. Wearables improve some labels while adding worker burden and safety review. Tactile sensors introduce drift, contamination, damage, replacement, and recalibration distributions. Aggressive privacy redaction removes context; weak redaction exposes workers and process knowledge. Independent factory-scale, long-duration validation, a common operator-minute denominator, and post-replacement replay remain scarce.
What to Learn Next
The next chapter moves to the point where human data and policy architecture meet: contact-rich work. Chapter 8 asks how to close the fourth loss defined here—force, touch, and contact state—through sensor synchronization, representation, force-aware policy design, and online correction. Human demonstration scale broadens candidate behavior, but release authority remains with real-robot contact traces and quality replay. In manufacturing manipulation, what the hand felt matters as much as what the camera saw.
References
- Chi, Cheng et al. (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv. Terry #35 · KO.
- Choi, Hojung et al. (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv. Terry #36 · KO.
- Xu, Mengda et al. (2025). DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation. arXiv. Terry #8 · KO.
- Fang, Hao-Shu et al. (2025). DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation. arXiv. Terry #10 · KO.
- Si, Zilin et al. (2025). ExoStart: Efficient learning for dexterous manipulation with sensorized exoskeleton demonstrations. arXiv. Terry #9 · KO.
- Qin, Yuzhe (2023). AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System. arXiv.
- Ding, Runyu (2024). Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning. arXiv.
- Kareer, Simar (2024). EgoMimic: Scaling Imitation Learning via Egocentric Video. arXiv.
- Zheng, Ruijie et al. (2026). EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv.
- Yang, Fengyu (2023). Touch and Go: Learning from Human-Collected Vision and Touch. arXiv.
- Feng, Ruoxuan (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-Tactile Sensors. arXiv.
- Choi, Hojung (2025). CoinFT: A Coin-Sized, Capacitive 6-Axis Force Torque Sensor for Robotic Applications. arXiv.
- Yu, Jiawen (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
- Huang, Jialei (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv.
- O'Neill, Abby (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
- Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
- AgiBot-World-Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
- Shaw, Kenneth (2023). LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. arXiv.
- Lambeta, Mike (2024). Digitizing Touch with an Artificial Multimodal Fingertip. arXiv.
- Figure AI (2026). Figure 03 + Helix 02: General-Purpose Humanoid System. Company product page.
- Physical Intelligence (2025). OpenPI: Open Source Robot Policy Stack. GitHub.
- Black, Kevin (2024). pi0: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
- Octo Model Team (2024). Octo: An Open-Source Generalist Robot Policy. arXiv.
- Mandlekar, Ajay et al. (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
- Sharma, Pratyusha et al. (2018). Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation. arXiv.
- Mandlekar, Ajay et al. (2018). RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation. arXiv.
- Wu, Philipp et al. (2023). GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators. arXiv.
- Zhao, Tony Z. et al. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
- Wang, Chen et al. (2024). DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation. arXiv.
- Zhaxizhuoma et al. (2024). FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset. arXiv.
- Ohkawa, Takehiko et al. (2026). YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale. Technical report.
- Xu, Chaoyi et al. (2026). RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning. arXiv.
- Lum, Tyler Ga Wei et al. (2025). Crossing the Human-Robot Embodiment Gap with Sim-to-Real RL using One Human Demonstration. arXiv.
- Fan, Jiacheng et al. (2026). RoboPaint: From Human Demonstration to Any Robot and Any View. arXiv. Terry #15 · KO.
- Fu, Zipeng et al. (2024). HumanPlus: Humanoid Shadowing and Imitation from Humans. arXiv. Terry #38.
- Hoque, Ryan et al. (2025). EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. arXiv. Terry #76 · KO.
- Sundaram, Subramanian et al. (2019). Learning the Signatures of the Human Grasp Using a Scalable Tactile Glove. Nature.
- Yin, Jessica et al. (2025). OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer. arXiv. Terry #18 · KO.
- Zhang, Chi et al. (2025). UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer. arXiv. Terry #16 · KO.
- Hoque, Ryan et al. (2021). ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning. arXiv.
- Dorsa Sadigh et al. (2017). Active Preference-Based Learning of Reward Functions. Robotics: Science and Systems.
- Roberto Calandra et al. (2018). More Than a Feeling: Learning to Grasp and Regrasp Using Vision and Touch. IEEE Robotics and Automation Letters.
- Runfa Blark Li et al. (2026). PhysGraph: Physically-Grounded Graph-Transformer Policies for Bimanual Dexterous Hand-Tool-Object Manipulation. arXiv preprint.
- Xingyu Peng et al. (2026). BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination. arXiv preprint.
- Alessio Palma et al. (2026). Bimanual Robot Manipulation via Multi-Agent In-Context Learning. arXiv preprint.
- Weiguang Zhao et al. (2026). Towards Robotic Dexterous Hand Intelligence: A Survey. arXiv preprint.
- OpenAI (2019). Learning Dexterous In-Hand Manipulation. International Journal of Robotics Research.