Chapter 3: Data Flywheels — Accumulating Learning Assets
Overview
A manufacturing data flywheel is not a warehouse that accumulates demonstrations. It is an operating system that returns evidence from human intent, robot-native failures, targeted correction, synthetic variation, and quality inspection to the next training and release decision. Its durable outputs are task definitions, calibrated scenes, failure taxonomies, replay sets, quality outcomes, and versioned lineage—not merely a larger folder of trajectories.
Recent collection systems widen the design space. GELLO and ALOHA make robot-matched teleoperation easier; Open-TeleVision extends immersive control to bimanual and whole-body behavior [25] [26] [27]. UMI, DexCap, DexUMI, and EgoDex seek learning signals without occupying the target robot [1] [29] [3] [31]. Intervention learning and synthetic trajectories concentrate new data around weak states and scene variations [32] [34] [37]. None automatically establishes contact truth, safe execution, or final product quality.
The loop closes only when those sources meet under one episode identity. Human demonstration records intent and recovery; robot rollout records the real cell; simulation and world models amplify selected variations; QA identifies failures that cause production loss. The central question is not whether teleoperation disappears. It is which costs move into tracking, retargeting, filtering, simulation calibration, supervision, and real validation—and whether the remaining real evidence covers deployment risk.
After reading this chapter... - Explain why humans, robots, simulation, and QA logs provide different labels. - Compare UMI, DROID, OXE, and AgiBot-style large collection through a manufacturing episode lens. - Argue why failure mining, replay sets, and rollback gates form the final link in the flywheel. - Design the collection sequence for a first production PoC.
Direct Teleoperation Preserves Robot-Native Signals
GELLO builds a low-cost leader with the kinematic structure of the target arm, using printed parts and commodity motors to make joint-space teleoperation intuitive [25]. The match simplifies pose correspondence, but a new arm family generally needs a new leader and calibration. ALOHA-style systems collect precise bimanual demonstrations through paired leader and follower arms, while Mobile ALOHA combines base motion with two-arm manipulation [26]. Their data aligns closely with robot execution, but collection occupies the robot and workcell and retains operator fatigue, reset, and safety-monitoring costs.
Open-TeleVision tracks head and hand motion and supplies active stereoscopic visual feedback for immersive humanoid teleoperation [27]. It can communicate high-dimensional whole-body and bimanual intent, yet immersion is not physical equivalence. Network latency, occlusion, robot joint limits, and collision guards alter the performed motion. A visually successful insertion does not expose achieved torque, tangential-force margin, internal material damage, or delayed inspection outcome unless the cell measures them separately.
Direct teleoperation records observations and intended actions on the target cameras, joints, gripper, controller, and clock. If force/torque and tactile sensing are installed, their traces can be synchronized to the demonstration. This alignment is valuable for precision insertion, coordinated hands, and tool use. It still produces candidate supervision rather than unquestionable ground truth: demonstrators vary, safety controllers modify commands, resets hide labor, and hardware drifts. Offline-imitation studies likewise show that performance depends on demonstration quality, diversity, observation modality, and learning choices, not just episode count [11].
| Collection route | Signal preserved well | Cost shifted or retained | Real-cell evidence still needed |
|---|---|---|---|
| GELLO/ALOHA-style direct control | Target-joint actions, cameras, control timing | Robot occupancy, leader build, fatigue, resets | Contact, long-horizon stability, tool wear |
| Open-TeleVision immersion | Whole-body intent, bimanual coordination, active view | Latency, occlusion, safety watch, retargeting | Collision margin, force limits, link failure |
| UMI/FastUMI portable capture | In-situ scenes and relative trajectories | Tracking, calibration, timing, feasibility filters | Executable motion and contact on target |
| Wearable/egocentric capture | Human-hand diversity and task context | Embodiment mapping, occlusion, annotation error | Touch, force, final product quality |
| Selective intervention | Policy-induced weak states and recovery | Watch time, takeover delay, reset, risk estimation | Missed failures and supervisor variance |
Robot-Free Capture Trades Occupancy for Conversion
UMI attaches cameras and tracking to a portable gripper, collects human manipulation in ordinary environments, and uses relative-pose actions plus latency matching to transfer data into a robot policy [1]. FastUMI decouples parts of the apparatus and uses off-the-shelf tracking to reduce deployment complexity [28]. The economic attraction is clear: a person can collect across sites without transporting the target robot, and the robot can remain available for production or high-value validation.
“Robot-free” does not mean “conversion-free.” Visual-inertial trajectories drift; relative actions must fit the target workspace and avoid singularities; fast human motion must respect controller rate, velocity, and acceleration limits. Texture-poor and reflective scenes challenge tracking. A similar gripper aperture does not guarantee the same arm clearance or approach direction. If feasibility filtering and replay are omitted, robot occupancy falls while unexecutable supervision rises.
DexCap combines portable localization with electromagnetic hand tracking, then connects capture to inverse kinematics and point-cloud imitation [29]. DexUMI constrains human motion through a wearable exoskeleton and visually replaces the human hand with the robot hand [3]. RealDexUMI uses the same lightweight dexterous hand, in-hand camera, and fingertip tactile module during wearable capture and deployment to reduce end-effector mismatch [30]. These designs narrow specific embodiment gaps; they do not equalize whole-arm dynamics, camera geometry, contact materials, or operating speed.
Robot-free collection can reduce target-robot occupancy while shifting cost into pose tracking, retargeting, action feasibility, temporal alignment, and real-robot validation. This conclusion is bounded by the cited tasks, embodiments, sensors, and evaluations; it does not establish universally lower total data cost or transfer without deployment-specific tests [28] [3] [30].
Egocentric corpora such as EgoDex can capture broad human-hand interactions and work context at much larger scale [31]. They also inherit occlusion, fast-motion annotation error, human-to-robot morphology gaps, and missing internal contact state. Video may be useful for visual and semantic pretraining without becoming measured action supervision for precision contact.
A Flywheel Is a Release Loop, Not a Pipeline
Many PoCs begin as a pipeline: collect demonstrations, train a policy, deploy the robot. Manufacturing needs more than that. Failures after deployment have to feed the next demonstration request, simulation perturbation, replay set, and release decision. Otherwise collection becomes a one-time annotation project and model updates separate from factory KPIs.
Figure 3.1. Factory data flywheel. The loop closes when human demonstrations, robot rollouts, simulation, and QA return through the same defect taxonomy. Source: author-created SVG.
| Flywheel source | Signal contributed | Manufacturing use | Bias to watch |
|---|---|---|---|
| Human demonstration | Intent, recovery, rare manipulation tricks | Initial coverage and recovery examples | Without force/contact, it stays close to video imitation |
| Robot rollout | Real actuators, fixtures, and cycle-time conditions | Production distribution and hardware drift | Safety limits restrict exploration of rare failures |
| Simulation | Asset variation, pose perturbation, guard tests | Pre-release regression and edge-case amplification | Tactile, fixture, and material models diverge from reality |
| QA/failure log | Defect code, rework, stop, operator takeover | Learning target and rollback criteria | Coarse labels block root-cause analysis |
Large corpora such as Open X-Embodiment and AgiBot World matter because they broaden policy priors [7] [8]. A manufacturer still has to read each source through its bias. Human video scales cheaply but is weak on contact force; robot rollouts match embodiment but are constrained in failure exploration.
Human Demonstrations Provide Intent and Recovery
The value of human demonstrations is not only the successful path; it is the reason for a correction and the way a person recovers from contact trouble. UMI opened a route for collecting teaching data without deploying the robot in the wild [1]. UMI-FT adds force signals so compliant manipulation does not lose as much contact information [2]. In a factory cell, that distinction is large. A visual trace may show that an insertion succeeded, but miss the moment when the operator felt binding and backed off.
Figure 3.2. UMI-FT adds force observability to in-the-wild demonstrations. Human demonstrations become manufacturing data when they preserve both intent and contact adjustment. Source: Choi et al. 2026, arXiv:2601.09988 Fig. 1.
DexUMI, DEXOP, and ExoStart reduce the cost of transferring dexterous human data into robot execution [3] [4] [5]. But a manufacturer should not stop at "the human hand did it." The paired evaluation has to measure which contacts disappeared on the target robot hand, which recoveries became impossible, and which defect codes appeared after transfer.
Robot-Native Rollouts Provide the Real Distribution
Robot-native data is valuable because it comes from the same actuators, cameras, tools, and fixtures as the deployed cell. DROID and Bridge Data show how real-world robot data can improve generalization [6] [10]. In manufacturing, that data also has to attach to lot, fixture, maintenance, and inspection state.
Robot rollouts reveal hardware drift. Gripper pads wear, suction cups collect residue, camera exposure changes, and fixtures move by small amounts. These changes are absent from many human demonstrations but responsible for a large share of production failures. A rollout episode should therefore include model version, hand or tool version, calibration version, shift, operator, and maintenance event.
Selective Intervention Collects Failure Boundaries
A policy trained from initial demonstrations compounds error outside its training distribution. Continuous human control is safe but hard to scale; unattended autonomy can create dangerous failures and costly resets. Selective intervention lets the policy act until it reaches an unfamiliar or risky state, then transfers control to a person and returns the correction segment to training. Its informational value comes from proximity to states the policy itself visits, not simply from segment length.
ThriftyDAgger learns a switching rule that requests supervisor control when novelty or estimated completion risk exceeds a budget-aware threshold [32]. Model-based runtime monitoring forecasts short latent futures and uses an intervention-trained failure classifier to request human attention [33]. Human-in-the-loop reinforcement learning combines demonstrations, a learned reward classifier, online interventions, resets, and controller engineering for precise tasks [34]. Together, these studies motivate concentrating supervision near failure boundaries rather than collecting only more clean successes.
Selective intervention targets policy-induced failure states, but its efficiency must count operator watch time, takeover latency, safety supervision, missed risk, and resets. This includes more than the duration of corrected segments. The cited results are bounded to limited tasks and supervision protocols; they do not establish the total cost or safety of continuous manufacturing oversight [32] [33] [34].
Intervention labels are not absolute truth. An expert may hear a vibration and intervene before a novice sees failure. Communication and safety-controller delays place the logged takeover after the underlying hazard. A monitor trained on previous interventions may miss a novel failure the policy has never visited. Episodes should retain supervisor identity, reason code, pre-takeover context, latency, safety-stop state, recovery, and reset—not just an action splice.
Simulation Perturbs the Distribution
Simulation is not a replacement factory. Its value is to perturb conditions that are rare, dangerous, or expensive to create in reality. RoboNet and related large-scale robot learning demonstrated the value of coverage across robots and scenes [9]. Simulation evaluation and Isaac Lab-style tooling provide mechanisms for testing policies before deployment [15] [16].
In manufacturing, simulation becomes more useful when it is tied to QA labels. Randomizing clearance, friction, and initial pose in an insertion task can be helpful, but synthetic success does not imply process success unless the randomized cases connect to real jam codes, scratch codes, and operator recoveries. Simulation perturbation should be designed around the replay failures that must be stressed before release.
Trajectories and World Models Amplify Real Seeds
Synthetic amplification includes at least three distinct routes. DemoGen edits a reconstructed 3D scene and adapts an action trajectory to produce spatial variants from a real demonstration [35]. DexMimicGen segments coordinated bimanual demonstrations and recombines them in simulation [36]. These methods are strongest when a known skill must cover new object placements or combinations. They do not automatically invent contact strategies absent from the seed. Reconstruction, segmentation, generation yield, and filtering labor belong in the data-cost denominator.
Physics simulation varies mass, friction, object pose, lighting, camera noise, fixture tolerance, and other conditions that are risky or slow to create in reality. Wider randomization is not inherently better. Variations unrelated to observed defects consume compute, while poorly calibrated contact can generate plausible but unexecutable success. A useful pipeline reports synthetic yield, rejection reasons, real replay correlation, and the defect family each variation is intended to cover.
Learned world models add another route: predict future video conditioned on an action, or infer a candidate action from an observed transition. DreamGen generates embodiment-conditioned video trajectories and derives pseudo-actions through an inverse-dynamics or latent-action model [37]. This may connect a small robot seed to broader visual conditions, but inverse dynamics is underdetermined. Multiple torque and contact histories can produce a similar visible result. Normal force, friction state, tactile distribution, tool vibration, and internal deformation are not uniquely recoverable from RGB frames.
The user-supplied Cosmos 3 action post-training video is first-party NVIDIA capability context for forward-dynamics, inverse-dynamics, and policy-mode interfaces [41]. It shows a tutorial workflow, not an independently replicated sample-efficiency result across tasks and embodiments. A demonstration count used in that workflow is therefore not generalized in this chapter. Generated action chunks and video futures are also not measured force, torque, or tactile supervision.
| Synthesis path | Directly observed | Inferred or missing | Real-cell qualification |
|---|---|---|---|
| Direct teleoperation | Robot observation and commanded action | Operator intent outside the interface | Replay on target tool, fixture, and safety controller |
| Wearable or egocentric video | Human task context and visible recovery | Robot action, contact force, executable timing | Retargeting checks plus limited robot reproduction |
| Inverse dynamics or world model | Visible state transition | Candidate action; hidden force and contact remain ambiguous | Feasibility filtering, contact measurement, and QA outcome |
| Physics simulation | Parameters and states encoded in the scene | Unmodeled wear, contamination, and material behavior | Parameter calibration and correlation with frozen real failures |
| Selective intervention | Policy-visited risk state and human correction | Unvisited hazards and supervision overhead | Latency audit, safety review, reset cost, and replay coverage |
The table is an observability map, not a ranking. Video-rich routes can widen context, simulation can stress controlled neighborhoods, and intervention can concentrate supervision near policy failures. None supplies every action, contact, and quality variable directly. The least replaceable real data is therefore the evidence that qualifies an inferred variable for the target cell.
Synthetic trajectories and learned video rollouts amplify a real seed distribution; they do not remove the need for real contact and outcome evidence. This claim remains bounded to the cited scene structures, tasks, inverse models, filtering procedures, and real evaluations, and must not be read as recovery of hidden force, torque, or tactile state from generated RGB [35] [36] [37].
RoboCat adapts to a new task from human demonstrations, generates successful robot experience, and returns selected experience to a later training mixture [38]. Its broader lesson is that “self-generated” data still needs a seed, success criterion, filtering, and mixture policy. If a policy repeatedly collects states it already handles, volume rises while weak coverage does not. Failure and uncertainty must remain assets rather than discarded runs.
NVIDIA's Physical-AI Strategy Is a Stack of Responsibilities
NVIDIA's strategy is best read as connected layers for scene authoring, physics and sensor execution, parallel learning, generative amplification, and robot-model training—not as one all-purpose simulator. OpenUSD provides composition and interchange for scene geometry, materials, joints, and semantics. Omniverse libraries and connectors bring those assets into established 3D authoring applications. The supplied Blender video is first-party context for preparing a “simulation-ready” scene; it is not independent evidence that a Blender scene has reality-equivalent physics [42].
NVIDIA's 2026 announcements describe embedding ovphysx, ovrtx, and scene-validation components into existing applications, with agents identifying missing physical properties and semantics [39] [40]. Structural completeness and parameter truth are different. A populated mass field may still contain the wrong mass. A plausible LiDAR preview does not establish calibration under temperature, reflective surfaces, contamination, or real sensor aging.
Isaac Sim executes robots, sensors, scenes, physics, and synthetic-data generation; Isaac Lab organizes parallel environments, tasks, randomization, training, and evaluation on top [43] [16]. Newton extends the intended physics layer toward contact and deformable-body computation. These components can generate candidate experience and run regression, but they do not identify process-specific friction or material behavior automatically and do not approve policy safety. Parameter identification, sensor calibration, and real-to-sim correlation remain separate work.
Cosmos belongs in the visual and behavioral generation layer. Action-conditioned models can propose visual futures or pseudo-actions; GR00T-family models belong in a later robot-policy training layer. Sharing an ecosystem does not collapse authoring, contact physics, action conversion, safety control, and quality inspection into one evidence claim.
OpenUSD/Omniverse authoring, Isaac simulation and learning, Cosmos augmentation, robot-policy training, and real-cell assurance address distinct gaps and must not be collapsed into one simulator. NVIDIA's official release and technical blog support only the announced functions and workflows, not independent factory-scale performance [39] [40]. Gao et al.[2] is a third-party survey of Isaac Sim; real deployment still requires calibrated physics, physical replay, quality records, and explicit safety and rollback gates.
| Gap | NVIDIA layer | Learning asset produced | What the layer does not establish |
|---|---|---|---|
| Scene interchange and authoring | OpenUSD and Omniverse libraries | Versioned scenes, semantics, reusable assets | Correct real mass, friction, and tolerance |
| Robot and sensor execution | Isaac Sim | Synthetic sensors and virtual rollouts | Real sensor aging, contact correlation, safety approval |
| Parallel learning and evaluation | Isaac Lab | Randomized tasks, training and regression results | Valid task definition or final quality acceptance |
| Contact and deformable physics | Newton and supported backends | Candidate contact and material behavior | Cell-specific identification and real equivalence |
| Visual and behavioral amplification | Cosmos family | Generated video, pseudo-actions, candidate futures | Hidden force/touch recovery or true failure probability |
| Robot foundation-policy training | GR00T family and data pipelines | Pretrained and post-trained policies | Cell control limits and quality responsibility |
| Deployment assurance | Real robot and manufacturing process | Frozen replay, calibration, quality, rollback evidence | An obligation transferable to the generative stack |
Collection-Order Walkthrough: A Tray-Picking PoC
The connector-insertion case in the Korean edition and the tray-picking case here are complementary localized examples. They use different contact and defect mechanisms to test the same sequence: failure taxonomy, robot-native evidence, synthetic stress, and QA-gated release. The shared lesson is to spend real robot time on variables that conversion or generation cannot qualify, rather than treating either cell as a universal template.
A tray-picking PoC looks different when it is designed as a flywheel. In the first week, the goal is not to train the largest robot policy. It is to watch how operators recover. When a missed pick happens, does the operator shake the tray, move the suction cup, pick a different edge, or pause because the ingredient has deformed? Those observations build the failure taxonomy and recovery vocabulary. Interfaces such as UMI and UMI-FT are useful here because they preserve intent and contact cues [1] [2].
In the second week, run limited robot-native rollouts. The goal is not maximum success; it is to reveal hardware failures that human demonstrations did not contain. Does suction pressure decay slowly? Does cup contamination increase missed picks? Does camera exposure vary by shift? The benefit of real-world robot data from DROID and Bridge Data becomes cell-specific telemetry at this stage [6] [10].
In the third week, use simulation as a failure amplifier. Perturb the pose and material conditions around real missed picks, crushes, and double picks to reject candidate policies. The question is not whether simulation success is high. It is whether candidates become unstable around real failure neighborhoods. QA closes the loop. Inspection rejects and operator takeovers must return to the source episodes, or the flywheel never closes.
Failure Coverage Instead of Source Quotas
Collection plans often say "1,000 human demos, 5,000 robot rollouts, 100,000 simulation trajectories." In manufacturing, those quotas should be rewritten by failure coverage. A missed pick may need human recovery, robot hardware drift, simulation pose perturbation, and QA reject evidence. A contamination failure may depend more on cleaning logs and inspection evidence.
| Failure family | Human source | Robot source | Simulation source | QA/release source |
|---|---|---|---|---|
| Missed pick | Recovery that moves the cup | Suction pressure and cup wear | Pose, lighting, material variation | Pick reject and retry count |
| Crush/deformation | Moment when the operator backs off | Force limit and gripper timing | Material-stiffness sweep | Downstream defect code |
| Wrong orientation | Regrasp cue | Jaw pose and release timing | Bin-pose perturbation | Orientation inspection |
| Line stop | Takeover reason | Safety stop and controller mode | Known unsafe-state replay | Rollback gate |
This table assigns roles to sources rather than counting source volume. Contact-aware policy work such as ForceVLA and AnyTouch is important for explaining contact failures [12] [14], but not every failure is a force problem. Some are inspection-taxonomy problems, some are hardware-drift problems, and some are operator-procedure problems. The flywheel becomes a data engine only when it keeps those causes separate.
A Factory Data Contract Is a Boundary of Responsibility
A manufacturing data flywheel must join human demonstrations, robot-native failures, correction, generation, and downstream quality labels under one episode identity. RoboNet, Bridge Data, and RoboCat demonstrate parts of this pattern at research scale, but closed mixtures and bounded tasks do not substitute for independent factory evidence linking defects, rework, safety stops, and release decisions [9] [10] [38].
The flywheel often breaks at the data key. Human demonstration files, robot bags, simulation rollouts, and QA spreadsheets can all use different IDs. When that happens, root-cause analysis collapses. The contract is therefore not just a schema; it is a boundary of responsibility. It defines who preserves raw sources, who creates features, who can block release, and who can execute rollback.
Figure 3.3. Factory data contract schema. If the episode key does not link sensors, controllers, QA, and release decisions, the flywheel does not close. Source: author-created SVG.
Offline demonstration work shows that demonstration quality and task diversity affect performance [11]. Factory deployment adds operating responsibility on top. Training episodes and release-gate episodes should be separated, and replay sets by defect code should be version-controlled artifacts. Only then can a model update be distinguished from overfitting to the latest collection round.
Contract Ownership: Who Can Stop Release?
The most important sentence in a factory data contract is who can stop deployment. A robot vendor may provide model updates, an integrator may install the workcell, and the manufacturer may carry quality responsibility. In that structure, data rights and release authority can drift apart. If the vendor sees telemetry but the manufacturer pays for defects, the flywheel's learning benefit and operating risk are separated.
The contract therefore needs more than raw data ownership. At minimum, the manufacturer should receive defect-code replay results, model-version diffs, rollback options, operator-takeover logs, and safety-stop summaries. Production-oriented companies such as Covariant, Dexterity, and Chef Robotics describe operating flywheels in product language [18] [19] [20]. If the manufacturer cannot execute release gates under its own QA taxonomy, the flywheel remains inside an external service.
How Simulation and QA Correct Each Other
Simulation has to return through the QA log. If simulated tray picking looks strong but real QA shows more contamination rejects, the simulation asset is missing a load-bearing failure mode. Conversely, if QA records a rare line stop that is risky to reproduce in the real cell, simulation can replay the neighborhood and reject unsafe release candidates. Without this mutual correction, simulation becomes a throughput metric and QA becomes a postmortem.
Isaac Lab and simulation-evaluation work provide useful tools for that correction [15] [16], but the process failure should decide which parameters to perturb. Clearance, friction, lighting, object pose, material stiffness, and fixture offset do not matter equally. If no one knows which parameter corresponds to a real defect, synthetic trajectory volume does not sharpen the learning target. A good flywheel lets QA choose the simulation questions.
Failure Mining Is Not the Last Step
Teams often treat failure mining as post-deployment analysis. In manufacturing, it is closer to the starting point. QA rejects, operator takeovers, line stops, safety stops, and rework tell the team what to collect next. Touch-and-Go and AnyTouch point to the role of contact signals in understanding failures [13] [12], while ForceVLA moves force-aware routing into the policy architecture [14].
Failure mining should produce three artifacts: the top failure codes and their cost, the sensor gap needed to explain each failure, and the replay items that block the next model release. Without those artifacts, a "data flywheel" is only a loop in name; in practice it becomes random accumulation.
Manufacturing Cell Checkpoint
Design the first flywheel by failure coverage, not by source quotas. For a cap-fastening PoC, start with cross-thread, under-torque, over-torque, cap scratch, and downstream leak as the taxonomy; then assign human demonstration, robot rollout, simulation perturbation, and QA evidence to each failure. For tray picking, missed pick, double pick, crush, contamination, and orientation error may be the organizing labels.
The second checkpoint is release authority. Who stops deployment when the new model fails a replay set? What diff and rollback option does the manufacturer receive when a robot vendor updates a cloud model? Production-oriented companies such as Covariant, Dexterity, and Chef Robotics show why operating flywheels matter [18] [19] [20], but the manufacturer's own defect taxonomy and data rights still need to be in the contract.
The third checkpoint is stale-data management. Human demonstrations age quickly after process changes, robot rollouts shift after maintenance, and simulation assets can lag behind fixture revisions. An episode needs the process version under which it was collected, not only the collection date. If stale data is not retired, the flywheel keeps rotating around an old factory.
Flywheel Cadence: Daily, Weekly, and Release Reviews
A flywheel is not a dashboard designed once and then ignored. Daily, weekly, and release-time reviews should ask different questions. Daily reviews watch line stops, safety stops, operator takeovers, and top defect codes. Weekly reviews open episode samples by defect and look for missing source coverage. Release reviews combine replay sets, simulation stress results, limited real rollout, and rollback plans.
| Review gate | Evidence entering the review | Decision or artifact |
|---|---|---|
| Daily operations | Stops, takeovers, resets, and current defect codes | Triage owner and episodes retained for diagnosis |
| Weekly coverage | Failure samples, sensor gaps, and source lineage | Next collection target and sources to freeze or reduce |
| Model release | Frozen replay, simulation stress, limited real rollout, and QA | Promote, hold, or roll back with named authority |
| Process or maintenance change | Tool, fixture, calibration, material, and inspection revisions | Retire stale evidence and rebuild affected replay cases |
Without that cadence, failure mining drifts into postmortem work. If missed picks rise on Monday and suction cups were changed on Wednesday, the team should inspect the hardware event before blaming Friday's model update. If hardware did not change but failures rise on one SKU, simulation perturbation or human recovery collection may need to be redesigned. Cadence creates a rhythm for conversation across data sources.
The flywheel meeting should ask, "Which data can we collect less of next week?" A failure that is already explained can be frozen into the replay set, while collection budget shifts to failures that remain unexplained. With that question, the data flywheel becomes a system for reducing collection cost rather than an infinite accumulation loop.
Another useful metric is failure half-life: the time from first detection of a defect to stable reduction through replay sets, simulation stress, robot rollout, and QA gates. A long half-life may mean the issue is not data volume, but a broken responsibility link. The human may know the recovery but it never transfers into robot rollout. Simulation may generate the failure but not match the QA label. A vendor update may arrive without replay evidence.
This metric translates the flywheel into operating language. A good flywheel is not the system that grows data volume forever; it is the system that shortens the time before the same failure stops recurring. The core of Chapter 3 is therefore feedback latency, not just source integration. Manufacturers should measure not only which data source is missing, but which link in the loop is slow.
The practical way to shorten feedback latency is to assign one owner to each persistent failure. A missed pick may belong to a process engineer, a force blind spot to a controls/data engineer, and an inspection mismatch to a quality engineer. Without ownership, flywheel meetings produce good observations but no next-week action. The more data a team has, the easier it is for responsibility to blur; the flywheel therefore needs an ownership board as much as a metric board. Owners should report not only whether the failure improved, but which source will be collected less because the cause is now explained.
That last clause matters operationally. If a cause is explained but the same high-rate log keeps growing, the flywheel has become storage cost rather than learning. Retiring or down-sampling a source is not a retreat from large data; it is evidence that the loop has converted one uncertainty into an engineering rule.
Open Questions and Failure Modes
Public research rarely exposes production telemetry. Human data is scaling faster than force and tactile observability. Simulation throughput can create false confidence when it is not tied to real QA labels. Vendor-owned data flywheels can externalize the manufacturer's process learning.
The dangerous state is high data volume with sparse failure coverage: many human demos without operator takeovers, many robot rollouts without defect codes, many simulation trajectories without real replay linkage. In that state, the model grows while the process does not learn.
Figure 3.4. Production data loop from observation through revalidation. This diagram is not a performance claim; it summarizes the operational chain that carries capability evidence through replay and quality decisions into release authority. Source: author-created SVG.
What to Learn Next
The flywheel specifies what to collect and how evidence returns, but it leaves a harder question: which virtual experience should be trusted for contact-rich work? Chapter 4, “Contact Models — Foundations for Control and Transfer,” examines contact models, controller boundaries, and sim-to-real validation. The episode lineage and frozen replay sets defined here become inputs to physical-parameter calibration and guarded policy tests there. The next question is not merely whether more data can be generated, but under which physical assumptions generated data predicts real contact and quality.
References
- Chi, Cheng et al. (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv. #35 Terry commentary.
- Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv.
- Xu, Mengda (2025). DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation. arXiv. #8 Terry commentary.
- Fang, Hao-Shu (2025). DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation. arXiv. #10 Terry commentary.
- Si, Zilin (2025). ExoStart: Efficient learning for dexterous manipulation with sensorized exoskeleton demonstrations. arXiv. #9 Terry commentary.
- Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
- O'Neill, Abby (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
- AgiBot-World Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
- Dasari, Sudeep (2019). RoboNet: Large-Scale Multi-Robot Learning. arXiv.
- Ebert, Frederik (2021). Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets. arXiv.
- Mandlekar, Ajay (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
- Feng, Ruoxuan (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-Tactile Sensors. arXiv.
- Yang, Fengyu (2022). Touch and Go: Learning from Human-Collected Vision and Touch. arXiv.
- Yu, Wenhao (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
- Li, Xingyu (2024). Evaluating Real-World Robot Manipulation Policies in Simulation. arXiv.
- Mittal, Mayank et al. (2025). Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning. arXiv.
- Toyota Research Institute (2024). Large Behavior Models for Robot Manipulation. Company technical post.
- Covariant (2024). RFM-1: Robotics Foundation Model. Company technical post.
- Dexterity (2025). Dexterity Foresight: AI Platform for Industrial Robot Workcells. Company product page.
- Chef Robotics (2025). ChefOS: AI Robotics Platform for Food Manufacturing. Company product page.
- Black, Kevin (2024). pi0: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
- Octo Model Team (2024). Octo: An Open-Source Generalist Robot Policy. arXiv. #55 Terry commentary.
- Brohan, Anthony (2022). RT-1: Robotics Transformer for Real-World Control at Scale. arXiv.
- Brohan, Anthony (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv.
- Wu, Philipp et al. (2023). GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators. CoRL.
- Fu, Zipeng et al. (2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv.
- Cheng, Xuxin et al. (2024). Open-TeleVision: Teleoperation with Immersive Active Visual Feedback. CoRL.
- Zhaxizhuoma et al. (2024). FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset. arXiv.
- Wang, Chen et al. (2024). DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation. RSS.
- Xu, Chaoyi et al. (2026). RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning. arXiv. #36 Terry related commentary.
- Hoque, Ryan et al. (2025). EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. arXiv. #76 Terry commentary.
- Hoque, Ryan et al. (2021). ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning. CoRL.
- Liu, Huihan et al. (2023). Model-Based Runtime Monitoring with Interactive Imitation Learning. arXiv.
- Luo, Jianlan et al. (2025). Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning. Science Robotics.
- Xue, Zhengrong et al. (2025). DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning. arXiv.
- Jiang, Zhenyu et al. (2024). DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning. arXiv.
- Jang, Joel et al. (2025). DreamGen: Unlocking Generalization in Robot Learning through Video World Models. arXiv.
- Bousmalis, Konstantinos et al. (2023). RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation. arXiv.
- NVIDIA (2026a). NVIDIA Agent Toolkit Expands With New Omniverse Libraries, Putting AI Agents to Work Building Simulation-Ready Worlds. Official release.
- NVIDIA (2026b). Integrate Physical AI Capabilities into Existing Apps with NVIDIA Omniverse Libraries. NVIDIA Technical Blog.
- NVIDIA (2026c). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official technical video.
- NVIDIA (2026d). Bringing Agent-Ready Simulation Into Blender. Official technical video.
- Gao, Sicong et al. (2026). NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics. Third-party survey, arXiv.
- Fan, Jiacheng et al. (2026). RoboPaint: From Human Demonstration to Any Robot and Any View. arXiv. #15 Terry commentary.
- Zheng, Ruijie et al. (2026). EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv.
- π0.6 Team (2025). π0.6: A Vision-Language-Action Model with Experience. Terry paper commentary [#4].
- ENPire Team (2026). ENPire: Robot Policy Self-Improvement. Terry paper commentary [#69].
- Jeffrey Mahler et al. (2017). Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics. Robotics: Science and Systems.
- Various (2024). DexH2R: Task-oriented Dexterous Manipulation from Human to Robots. arXiv preprint.
- Sungjae Park et al. (2025). Learning to Transfer Human Hand Skills for Robot Manipulations. arXiv preprint.
- Marc Peter Deisenroth et al. (2011). PILCO: A Model-Based and Data-Efficient Approach to Policy Search. ICML / Artificial Intelligence.
- Ankur Handa et al. (2020). DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System. ICRA 2020.
- Jeannette Bohg et al. (2014). Data-Driven Grasp Synthesis: A Survey. IEEE Transactions on Robotics.
- David J. Montana (1988). The Kinematics of Contact and Grasp. International Journal of Robotics Research.