Chapter 1: Data Scale — Conditions for Manufacturing Coverage
Overview
In manufacturing manipulation, large data does not simply mean more demonstration videos. A pick, insertion, or wiping motion that looks identical on camera can belong to a different distribution when the SKU, fixture, surface condition, tool wear, operator intervention, or inspection rule changes. Open X-Embodiment and DROID show how robot data can broaden embodiment and task families [1] [2], but a factory cell only gains an improvement loop when each attempt is tied to quality outcomes and operating logs.
The core unit in this chapter is not a successful motion; it is an inspectable attempt episode. An episode has to preserve observation, action, contact state, controller mode, operator override, defect code, rework outcome, and rollback condition together. Bridge Data and RoboNet demonstrate the value of cross-domain robot data [3] [4]. In manufacturing, however, the decisive question is not only what row entered a dataset, but which process state produced it and which quality result ended it.
After reading this chapter... - Explain manufacturing manipulation as an episode-to-quality linkage problem rather than a model-score contest. - Argue why failure distribution, reproducibility, and replay sets come before headline success rate. - Distinguish vision-only data from force- and tactile-rich data for factory tasks. - Define the minimum logs a manufacturer should own and the export fields it should require from a vendor.
Why Factory Manipulation Is Stricter Than Generic Robot Data
General robot datasets mix robots, scenes, and instructions to improve policy extrapolation. Factory cells do almost the opposite: they expose small differences inside repetitive work where small differences can become defects. In cap fastening, the thread start, tiny bottle deformation, torque limit, and leak-test code change the meaning of an episode. In food-tray picking, ingredient moisture, tray-pocket contamination, suction-cup wear, and line speed can turn the same action into a different outcome.
RoboNet pooled action-conditioned videos across institutions and robot arms [4]. MIME paired third-person human videos with kinesthetic robot trajectories for human-to-robot imitation [25]. Both were important steps beyond a single laboratory and task, but their reported settings remain dominated by short-horizon interactions or household tasks, and paired human and robot trajectories are not force-equivalent. The frontier RoboTacDex dataset synchronizes multiview RGB/depth, touch, robot state, and semantic labels, yet remains bounded to one humanoid platform, a moderate task set, and a forthcoming public release [26].
Dataset scale is therefore not sufficient evidence of factory coverage unless every episode retains failure and outcome context [4] [25] [26]. Coverage is not simply a task count. It is the measured region formed by SKU and lot, fixture revision, material state, approach, contact phase, inspection rule, and cycle-time limit. A million successful clips provide no coverage for a thread-damage or seal-leak failure that was never labeled and cannot be replayed. This is a manufacturing acceptance criterion synthesized here; the cited datasets establish cross-robot or multimodal research coverage, not independently audited factory coverage.
Figure 1.1. Manufacturing data flywheel. Data becomes an accumulating asset only when production episodes, failure logs, replay sets, and model updates form a closed loop. Source: author-created SVG.
Large models are one part of that loop. RT-1, RT-2, and pi0-style systems provide broad task priors [7] [8] [15], but a manufacturer still has to trace whether those priors reduce defect codes in a specific cell. S10 therefore asks less "which model is largest?" and more "which attempt remains explainable enough to drive the next improvement?"
Historical Scaling: What Large Collection Did and Did Not Establish
Classical manipulation made geometry, dynamics, contact constraints, planners, and controllers explicit. That lineage remains valuable because it supplies interpretable force and safety limits. Learning from demonstration then separated demonstration acquisition, human-robot correspondence, policy derivation, and policy improvement as distinct problems [28]. Large self-supervised collection added a different capability: robots could repeat attempts, preserve failures, and fit visual controllers where friction, deformation, occlusion, and grasp outcome were difficult to model completely.
Large real-robot collection established an important scaling lineage, but every reported gain remains bounded to its robot, task distribution, and evaluation protocol [5] [27] [6]. Levine et al. learned closed-loop visual grasping from experience gathered across multiple robots, while QT-Opt combined off-policy reinforcement learning with distributed collection at larger scale [6] [5]. These results show that repeated interaction can improve grasping under the reported settings; they do not by themselves establish transfer to insertion, assembly, wiping, deformable handling, or downstream quality.
The well-known 50,000-try and 700-robot-hour grasping result is a bounded 2016 collection example, not a universal production-data requirement [27]. Its binary grasp label supports the question of whether an object was lifted, not whether it was placed in the correct downstream pose, damaged, processed within takt time, or accepted by inspection. A manufacturer should remeasure which variation and failure modes move its learning curve rather than copy the interaction count into a data budget.
The historical synthesis is not “collect more and discard models.” Requested policy actions, safety-controller corrections, sent commands, and measured motion still need distinct records. Without that separation, a policy error can be credited as a controller success, or a safety projection can be mistaken for evidence that the learned action was valid.
The Minimum Schema for a Factory Episode
For a first PoC, the episode schema should be narrower and harder-edged than a research demonstration format. Cameras, robot commands, force sensors, inspection stations, and operator interfaces commonly use different clocks and buffers. A start and end timestamp cannot reconstruct the rare event if clock offsets, missing intervals, and asynchronous inspection are not recorded.
A manufacturing episode must connect requested and executed actions to intervention, reset, result, defect, and calibration/version state [29] [30] [26]. SOTER combines an uncertified high-performance controller with a certified lower-performance safety controller under a safety specification [30]. Model-based runtime monitoring with interactive imitation learning forecasts risky futures and requests selective human monitoring from intervention-trained labels [29]. Both lines show why a policy proposal and the action admitted by the factory control stack cannot be stored as one field. This separation is a manufacturing logging prescription synthesized from runtime-assurance and intervention evidence; the cited drone and research-manipulation studies do not validate the full episode schema or confer factory safety assurance.
The table below lists fields that are hard to omit for repeated manual work such as cap fastening, tray picking, and press-fit insertion.
| Episode field | Why it matters | Failure when missing |
|---|---|---|
| SKU, lot, fixture revision | Separates distributions that look like the same task | A lot-specific failure is misread as a model error |
| Hand/tool ID and calibration | Tracks grip geometry and sensor drift | Performance drops after a hardware change become unexplained |
| Proposed, sent, and executed action | Separates policy, guard, and actuator effects | A safety correction is counted as policy competence |
| Force/torque or tactile event | Reveals slip, jam, and over-force | Visual success hides physical defects |
| Intervention, reset, and recovery code | Preserves labor and failure-boundary evidence | Increased human work is counted as autonomy |
| Controller mode and guard state | Separates learned action from deterministic fallback | A policy update increases safety stops without explanation |
| Inspection result and defect code | Ties learning to process KPIs | Success clips grow while scrap and rework stay flat |
| Data, model, and calibration versions | Enable reproduction and rollback | No one can explain which update changed the cell |
| Replay-set membership | Creates regression tests for the next release | The same failure reappears across versions |
This schema is also a lock-in question. If data rights stay entirely with the robot provider, the manufacturer cannot verify why a model improved. Claims from production-oriented companies such as Covariant, Dexterity, and Chef Robotics matter [16] [17] [18], but if the manufacturer cannot export episodes and release-gate results, the cause of improvement remains inside an external black box.
Failure Distribution Before Success Rate
Average success rate is a late metric in a factory cell. Early on, the useful work is to separate repeatable failures, sensor blind spots, and hardware drift. QT-Opt and large-scale grasping showed the power of repeated collection for vision-based manipulation [5] [6], but in a factory the target for recollection stays vague unless failures are tied to inspection codes.
The reason is rollback. A new policy that raises aggregate success by two points is not deployable if it increases scratch defects on one SKU. Conversely, a policy with similar average success may be valuable if it removes a rare failure that causes line stops. A replay set is therefore not just a benchmark; it is a deployment guard. Every production release should replay recent safety stops, over-force events, missed picks, bad torque cases, and operator takeovers.
Contact Signals Are Not Optional
Some work is well served by cameras alone. Insertion, wiping, food handling, cable routing, and deformable pouch handling are different: contact state often determines success. DIGIT-style tactile fingertips show how contact can become an image-like field [14], while AnyTouch and ForceVLA point toward policies that consume tactile and force signals directly [12] [24].
Pose and RGB alone can under-specify contact-rich execution because force, touch, compliance, and downstream quality can differ for visually similar trajectories [26] [33] [31]. RoboTacDex adds synchronized touch and semantic labels to multiview vision and robot state [26], while DexForce converts fingertip force/torque recorded during kinesthetic teaching into force-informed action targets [33]. Neither establishes universal factory transfer: RoboTacDex is platform- and release-bounded, and DexForce uses one hand platform, two force-torque sensors, per-task compliance tuning, and a quasi-static spring-model assumption.
Figure 1.2. Digit 360 tactile fingertip as contact-rich data source. Contact location, pressure pattern, and slip changes are difficult to recover from vision logs alone. Source: Meta AI Research Digit 360 GitHub, MIT License.
For manufacturing, tactile sensing is valuable less because it makes a hand look more human and more because it improves failure attribution. It records whether the robot pressed too hard, touched the fixture first, rotated the part, or slipped because of surface contamination. Without that information, the same failure can be misassigned to hardware, controller, policy, or process conditions.
Decision Walkthrough: A Cap-Fastening Cell
Imagine a first cap-fastening cell. A human places a cap on the bottle neck, presses lightly to find the thread start, rotates while feeling for binding, and backs off immediately when torque feels wrong. In video, that process looks like "grasp and turn." As manufacturing data, it separates thread start, axial force, torque slope, cap tilt, and leak-test outcome. Without that separation, a policy learns the visible rotation but not the moment that creates cross-thread damage.
The first experiment is not to choose a larger foundation model. It is to record 200-300 manual cycles under the target episode schema and discover which defects carry cost. If under-torque dominates, torque trace and final inspection matter most. If scratches dominate, contact path and gripping force matter more. If leaks are discovered downstream, the episode ID must survive into that later inspection station. Broad priors from DROID or Open X-Embodiment can help [1] [2], but the learning target comes from the cell's defect taxonomy.
The second experiment is to turn robot failures into a replay set. If the early policy creates cross-thread failures, those episodes are not just logs; they become release-blocking test cases. The next model has to reduce the failure under the same SKU, fixture, and torque limit. A model that improves headline success while increasing operator takeovers or cycle-time variance is not a manufacturing improvement. The episode therefore carries operating metrics, not only success labels.
When the Data Lake Collapses the Episode
A common failure is to push every sensor stream into a data lake and assume alignment can be solved later. Timestamp alignment alone is not enough. Camera frames, robot commands, force events, inspection results, and operator notes often use different clocks and buffers. Rare failures are especially hard to reconstruct because safety stops and operator takeovers happen exactly when logs become messy.
| Where data collapses | Factory symptom | Collection design response |
|---|---|---|
| Sensor clock mismatch | Contact event and image frame do not line up | Store episode boundaries and hardware clock offsets |
| Delayed inspection | A later defect detaches from the learning episode | Route downstream QA results back to the original episode ID |
| Free-text operator notes | One failure appears under many names | Store both free text and standardized defect code |
| Missing release record | No one can explain which update improved the cell | Freeze model, dataset, and replay-set hashes in the release record |
This table may look like data engineering detail, but it is a condition for learning. If the cause of failure is not attached to the episode, the next collection round becomes "collect more" rather than "collect the missing signal." Cross-domain generalization from Bridge Data and RoboNet is valuable [3] [4], but manufacturing needs both wider priors and narrower defect attribution.
Growing From a Small Cell to a Large-Data Strategy
Going straight from a first cell to a general policy deployment usually blurs why data is needed. A safer sequence is narrow cell, strong logging, bounded policy, and release replay. First stabilize the episode schema. Then build the failure taxonomy. Only then import broader priors, and do not deploy them if they fail the cell-specific replay set.
The same sequence helps read company strategies. Covariant and Dexterity use the language of production workcells and operating data loops [16] [17]. Chef Robotics targets food manufacturing, a domain with both repetition and variation [18]. The manufacturer should look past vocabulary and ask for exportable evidence: which defect changed, which data source explained the change, and how rollback is executed.
Human Demonstrations Need Force Awareness
UMI and DROID show that human-centered collection interfaces can ease the robot data bottleneck [10] [2]. Yet manufacturing does not close the loop by collecting many videos of humans doing the task well. The useful part of human skill is often a forceful or compliant adjustment: twisting slightly during insertion, backing off at the moment of jamming, or regrasping just before slip. If those signals are not recorded, the robot receives only the visible shell of the skill.
Figure 1.3. Force-informed demonstrations convert contact into learnable action. For contact-rich work, human demonstrations must include force and compliance to become learning units. Source: Chen et al. 2025, arXiv:2501.10356 Fig. 1.
Interfaces such as UMI-FT reduce that gap by adding force observability [11]. A manufacturer should therefore evaluate demonstrations by the contact events they preserve, not only by their count. Operator correction is especially important. A takeover timestamp, reason, and recovery motion are high-value labels for regions the next policy should avoid or handle differently.
Evidence Tiers and Manufacturing Readiness
S10 does not read all references with the same confidence. Papers are useful for methods and experimental conditions. Company technical pages reveal product direction and public operating claims. Press-like announcements belong in a watchlist unless the claim is backed by primary evidence. Toyota Research Institute's large behavior model work shows the direction of faster robot teaching [19], but factory deployment still depends on cell-level QA, operator overrides, rollback authority, and data export.
Historical and frontier sources answer different questions. Large real-robot studies establish the value of repeated collection under their reported settings. Multi-institution datasets test broader distributional support. Force and tactile datasets expose contact variables that vision may omit, while runtime monitoring turns interventions into labels about risky states. None of these evidence types alone establishes production readiness, and a paper result, first-party technical announcement, product demonstration, and independently audited deployment should not be treated as interchangeable.
| Evidence source | What it directly supports | What it does not directly support | Manufacturing interpretation |
|---|---|---|---|
| Large real-robot collection | Benefits of scaling under a specified robot and task distribution | A universal interaction budget for every process | Remeasure the learning curve in each cell |
| Multi-institution or multi-task data | Broader priors and the possibility of transfer | Improvement in a named defect or cycle-time KPI | Separate pretraining evidence from cell acceptance |
| Force and tactile data | Added observability of contact state and force intent | Sensor durability or production reliability | Test whether the modality separates a costly failure |
| Runtime monitoring and intervention data | Labels for risky states and recovery behavior | Guaranteed detection of every novel failure | Record missed interventions and supervisor latency |
There is a substantive disagreement about broad priors. One position is that broad pretraining can reduce the number of local demonstrations needed for adaptation; RT-1 and Open X-Embodiment provide evidence with which to test that hypothesis [7] [1]. The counterposition is that embodiment, action-space, and contact-modality differences can make a common representation discard details that control factory quality. Current evidence does not justify turning either position into a universal law. A report should separately state whether a broad prior reduced initial exploration and how many real attempts were still needed to close local defect modes. Cross-robot task coverage is not contact equivalence, and simulation or visual diversity is not an audited downstream quality result.
This distinction matters in buy/build decisions. A claim may be "large-data driven" while referring to a benchmark score, a teleoperation dataset, or a production-cell defect reduction. Those are different evidence tiers. When watching a vendor demo, manufacturers should ask four questions: can failure episodes be exported under our defect taxonomy; can replay results be compared before and after a model update; who owns force, tactile, and raw video retention; and can the site lead execute rollback without waiting for the vendor?
Translating Evidence Tiers Into Operating Metrics
Task success in a paper is useful for comparing methods, but it needs translation before it becomes a factory purchase criterion. RT-1, RT-2, and pi0-style systems provide broad priors for instructions, objects, and situations [7] [8] [15]. In a cap-fastening cell, that evidence must be reread through thread damage, leak rejects, operator retightening, and safety stops. The same model may be helpful in one cell and too general for another.
Company technical pages need the same translation. A phrase such as production ready can imply deployment maturity, but public pages rarely expose full defect distributions and rollback authority. A manufacturer should map each source to an internal metric. Which hypothesis does a paper result support? Which failure mode is hidden by a demo? Which export and support promises appear in the official product material? References become useful procurement evidence only when they are translated into operating risk.
Buy/Build Memo: Contracting for Data Ownership
After the first technical review, the buy/build decision should move from a model scorecard to a data-ownership memo. Building internally gives the manufacturer direct ownership of raw episodes and replay sets, but raises the cost of models and operations tooling. Buying a vendor solution can speed deployment and support, but failure attribution and model-update explanation may remain outside the factory. That tradeoff has to become contract language, not only strategy language.
The memo needs at least five items. First, retention period and export format for raw video, force, tactile, robot command, and QA labels. Second, dataset diffs and replay results for every model update. Third, cause codes for operator takeovers and safety stops. Fourth, the audit scope for vendor-derived features. Fifth, authority for emergency rollback and long-horizon rollback. Without those items, the manufacturer contributes data without owning the process learning.
| Decision item | Signal for internal build | Signal for an external solution | Evidence retained either way |
|---|---|---|---|
| Raw data and lineage | Process knowledge and sensor originals require direct control | Complete episodes can be exported in a contracted format | Schema, calibration versions, retention and deletion record |
| Updates and regression | An internal team can operate training and evaluation versions | The provider supplies update diffs and replay results | Data/model hashes and defect-specific before/after results |
| Operations and safety | The site integrates controllers, monitors, and recovery | Provider scope is separated from site safety authority | Intervention/stop causes and emergency rollback procedure |
| Exit and transfer | Schemas and tests will be reused in the next cell | Data and operating knowledge transfer at contract exit | Exit checklist, unresolved failures, reproducible test bundle |
Privacy and process IP belong in the same memo. Worker video contains bodies and shop-floor know-how. Process video exposes fixtures and sequences. A large-data strategy is therefore also a governance strategy. Some data should remain on-premise, some should be shared only as features, and some should be deleted after learning. If those rules are not set early, a successful PoC can create a larger legal and security problem.
The last item is the exit criterion. At the end of a PoC, the manufacturer should not write only "success rate improved." It should record which defect fell, which log explained the cause, and which schema can be reused in the next cell. If the exit criterion is vague, pilots multiply without leaving manufacturing learning. If defect-specific improvement and an exportable episode schema remain, a small cell becomes a data standard for the next cell.
This criterion connects research and products. Papers show possible methods, vendors provide implementation and support, and manufacturers know process loss and responsibility. The episode is the unit that binds those layers. The large-data problem in this chapter is therefore not only scale; it is deciding which attempts become part of the company's operating memory.
Manufacturing Cell Checkpoint
The first PoC should be the task where the log can close, not the flashiest task. For cap fastening, attach torque trace, thread start, leak test, and operator retightening to the episode. For tray picking, attach suction pressure, missed-pick recovery, ingredient deformation, and inspection reject. Without those fields, success videos may grow while process KPIs stay flat.
Review the data pipeline in three steps. First, confirm that the episode ID links every sensor, controller, and inspection record. Next, confirm that failure codes populate replay sets. Finally, confirm that model releases connect to field rollback. If any of those links breaks, the large-data strategy becomes external demo consumption rather than a manufacturing asset.
Open Questions and Failure Modes
Worker video and process IP are learning assets, but also privacy and trade-secret risks. Tactile and force sensors introduce calibration, cleaning, and durability costs. Sim-to-real throughput does not guarantee factory performance unless it is tied to real QA labels. Company claims often expose too little production telemetry for independent comparison.
A second debate is how aggressively to collect failure data. Failures expose policy boundaries, but they can damage equipment and create scrap. ThriftyDAgger requests supervisor control in novel or high-risk states to improve data efficiency, but it depends on risk-model calibration and continued supervisor availability [34]. The operating rule should therefore specify which near failures and interventions may be logged inside the safety envelope, and how genuinely dangerous attempts are blocked, rather than simply calling for more failures.
The most common failure mode is "large data without causal explanation." If a missed pick cannot be attributed to hand wear, SKU change, lighting drift, or policy regression, the next data collection round repeats the same confusion. In S10, large-data manipulation means building the operating design that improves that attribution.
Figure 1.4. Manufacturing release gate from data coverage to update approval. This diagram is not a performance claim; it summarizes the operational chain that carries capability evidence through replay and quality decisions into release authority. Source: author-created SVG.
What to Learn Next
Chapter 2 applies the episode view to end-effector choice. Some cells are better served by suction or two fingers; others cannot observe the cause of failure without dexterous hands and tactile or force-rich sensing. The comparison asks how each end effector changes state space, collection cost, contact observability, and acceptance testing rather than assuming that more degrees of freedom are always better.
References
- O'Neill, Abby (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
- Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
- Ebert, Frederik (2021). Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets. arXiv.
- Dasari, Sudeep (2019). RoboNet: Large-Scale Multi-Robot Learning. arXiv.
- Kalashnikov, Dmitry (2018). QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation. arXiv.
- Levine, Sergey (2016). Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection. arXiv.
- Brohan, Anthony (2022). RT-1: Robotics Transformer for Real-World Control at Scale. arXiv.
- Brohan, Anthony (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv.
- AgiBot-World Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
- Ha, Huy (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv.
- Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv.
- Feng, Ruoxuan (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-Tactile Sensors. arXiv.
- Shaw, Kenneth (2023). LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. arXiv.
- Lambeta, Mike (2020). DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation. arXiv.
- Black, Kevin (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
- Covariant (2024). RFM-1: Robotics Foundation Model. Company technical post.
- Dexterity (2025). Dexterity Foresight: AI Platform for Industrial Robot Workcells. Company product page.
- Chef Robotics (2025). ChefOS: AI Robotics Platform for Food Manufacturing. Company product page.
- Toyota Research Institute (2024). Large Behavior Models for Robot Manipulation. Company technical post.
- NVIDIA (2025). Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning. NVIDIA Research.
- Mandlekar, Ajay (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
- Chi, Cheng (2023). Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv.
- Zhao, Tony Z. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
- Yu, Wenhao (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
- Sharma, Pratyusha et al. (2018). Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation. CoRL.
- Wang, Xinyi et al. (2026). RoboTacDex: A Dexterous Visual-Tactile-Action Dataset for Humanoid Manipulation. arXiv.
- Pinto, Lerrel and Gupta, Abhinav (2016). Supersizing Self-Supervision: Learning to Grasp from 50K Tries and 700 Robot Hours. ICRA.
- Argall, Brenna D. et al. (2009). A Survey of Robot Learning from Demonstration. Robotics and Autonomous Systems.
- Liu, Huihan et al. (2023). Model-Based Runtime Monitoring with Interactive Imitation Learning. arXiv.
- Desai, Ankush et al. (2019). SOTER: A Runtime Assurance Framework for Programming Safe Robotics Systems. DSN.
- Billard, Aude and Kragic, Danica (2019). Trends and Challenges in Robot Manipulation. Science.
- Fang, Hao-Shu et al. (2023). RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot. arXiv.
- Chen, Claire et al. (2025). DexForce: Extracting Force-Informed Actions from Kinesthetic Demonstrations for Dexterous Manipulation. RA-L. 한국어 해설 · English note.
- Hoque, Ryan et al. (2021). ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning. CoRL.
- Corey Lynch et al. (2020). Learning Latent Plans from Play. CoRL.
- Tony Z. Zhao et al. (2024). ALOHA Unleashed: A Simple Recipe for Robot Dexterity. Conference on Robot Learning (CoRL).
- Kailin Li et al. (2025). ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025.
- Roland S. Johansson et al. (2009). Coding and Use of Tactile Signals from the Fingertips in Object Manipulation Tasks. Nature Reviews Neuroscience.
- Chelsea Finn et al. (2017). Deep Visual Foresight for Planning Robot Motion. IEEE ICRA.
- Luis Sentis et al. (2005). Synthesis of Whole-Body Behaviors through Hierarchical Control of Behavioral Primitives. International Journal of Humanoid Robotics.
- Matthew T. Mason (1986). Mechanics and Planning of Manipulator Pushing Operations. International Journal of Robotics Research.
- Wenzhen Yuan et al. (2017). GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force. Sensors (MDPI).