Consolidated References
277 references
[1] O'Neill, Abby (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
[2] Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
[3] Ebert, Frederik (2021). Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets. arXiv.
[4] Dasari, Sudeep (2019). RoboNet: Large-Scale Multi-Robot Learning. arXiv.
[5] Kalashnikov, Dmitry (2018). QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation. arXiv.
[6] Levine, Sergey (2016). Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection. arXiv.
[7] Brohan, Anthony (2022). RT-1: Robotics Transformer for Real-World Control at Scale. arXiv.
[8] Brohan, Anthony (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv.
[9] AgiBot-World Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
[10] Ha, Huy (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv.
[11] Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv.
[12] Feng, Ruoxuan (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-Tactile Sensors. arXiv.
[13] Shaw, Kenneth (2023). LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. arXiv.
[14] Lambeta, Mike (2020). DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation. arXiv.
[15] Black, Kevin (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
[16] Covariant (2024). RFM-1: Robotics Foundation Model. Company technical post.
[17] Dexterity (2025). Dexterity Foresight: AI Platform for Industrial Robot Workcells. Company product page.
[18] Chef Robotics (2025). ChefOS: AI Robotics Platform for Food Manufacturing. Company product page.
[19] Toyota Research Institute (2024). Large Behavior Models for Robot Manipulation. Company technical post.
[20] NVIDIA (2025). Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning. NVIDIA Research.
[21] Mandlekar, Ajay (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
[22] Chi, Cheng (2023). Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv.
[23] Zhao, Tony Z. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
[24] Yu, Wenhao (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
[25] Sharma, Pratyusha et al. (2018). Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation. CoRL.
[26] Wang, Xinyi et al. (2026). RoboTacDex: A Dexterous Visual-Tactile-Action Dataset for Humanoid Manipulation. arXiv.
[27] Pinto, Lerrel and Gupta, Abhinav (2016). Supersizing Self-Supervision: Learning to Grasp from 50K Tries and 700 Robot Hours. ICRA.
[28] Argall, Brenna D. et al. (2009). A Survey of Robot Learning from Demonstration. Robotics and Autonomous Systems.
[29] Liu, Huihan et al. (2023). Model-Based Runtime Monitoring with Interactive Imitation Learning. arXiv.
[30] Desai, Ankush et al. (2019). SOTER: A Runtime Assurance Framework for Programming Safe Robotics Systems. DSN.
[31] Billard, Aude and Kragic, Danica (2019). Trends and Challenges in Robot Manipulation. Science.
[32] Fang, Hao-Shu et al. (2023). RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot. arXiv.
[33] Chen, Claire et al. (2025). DexForce: Extracting Force-Informed Actions from Kinesthetic Demonstrations for Dexterous Manipulation. RA-L. 한국어 해설 · English note.
[34] Hoque, Ryan et al. (2021). ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning. CoRL.
[35] Corey Lynch et al. (2020). Learning Latent Plans from Play. CoRL.
[36] Tony Z. Zhao et al. (2024). ALOHA Unleashed: A Simple Recipe for Robot Dexterity. Conference on Robot Learning (CoRL).
[37] Kailin Li et al. (2025). ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025.
[38] Roland S. Johansson et al. (2009). Coding and Use of Tactile Signals from the Fingertips in Object Manipulation Tasks. Nature Reviews Neuroscience.
[39] Chelsea Finn et al. (2017). Deep Visual Foresight for Planning Robot Motion. IEEE ICRA.
[40] Luis Sentis et al. (2005). Synthesis of Whole-Body Behaviors through Hierarchical Control of Behavioral Primitives. International Journal of Humanoid Robotics.
[41] Matthew T. Mason (1986). Mechanics and Planning of Manipulator Pushing Operations. International Journal of Robotics Research.
[42] Wenzhen Yuan et al. (2017). GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force. Sensors (MDPI).
[43] Lambeta, Mike (2024). Digitizing Touch with an Artificial Multimodal Fingertip. arXiv.
[44] Bhirangi, Raunaq (2021). ReSkin: versatile, replaceable, lasting tactile skins. arXiv.
[45] Bhirangi, Raunaq (2024). AnySkin: Plug-and-play Skin Sensing for Robotic Touch. arXiv.
[46] Choi, Hojung (2025). CoinFT: A Coin-Sized, Capacitive 6-Axis Force Torque Sensor for Robotic Applications. arXiv.
[47] Xu, Mengda (2025). DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation. arXiv. #8
[48] Fang, Hao-Shu (2025). DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation. arXiv. #10
[49] Si, Zilin (2025). ExoStart: Efficient learning for dexterous manipulation with sensorized exoskeleton demonstrations. arXiv.
[50] Huang, Yuhang (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv.
[51] Hao, Yaru (2025). TLA: Tactile-Language-Action Model for Contact-Rich Manipulation. arXiv.
[52] Figure AI (2026). Figure 03 + Helix 02: General-Purpose Humanoid System. Company product page.
[53] Physical Intelligence (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. Company research post.
[54] Bicchi, Antonio (2000). Hands for Dexterous Manipulation and Robust Grasping: A Difficult Road Toward Simplicity. IEEE Transactions on Robotics and Automation.
[55] Capsi-Morales, Patricia et al. (2020). Exploring the Role of Palm Concavity and Adaptability in Soft Synergistic Robotic Hands. IEEE Robotics and Automation Letters.
[56] Agarwal, Arpit et al. (2025). A Modularized Design Approach for GelSight Family of Vision-based Tactile Sensors. The International Journal of Robotics Research.
[57] Fang, Bin et al. (2025). Force Measurement Technology of Vision-Based Tactile Sensor. Advanced Intelligent Systems.
[58] Gupta, Harsh et al. (2025). Sensor-Invariant Tactile Representation. arXiv.
[59] Zhao, Zihang et al. (2025). Embedding High-Resolution Touch across Robotic Hands Enables Adaptive Human-Like Grasping. Nature Machine Intelligence.
[60] Almeida et al. (2025). The Role of Touch: Towards Optimal Tactile Sensing Distribution in Anthropomorphic Hands for Dexterous In-Hand Manipulation. arXiv.
[61] Zhang, Chi et al. (2025). UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer. arXiv.
[62] Felix Berkenkamp et al. (2017). Safe Model-based Reinforcement Learning with Stability Guarantees. NeurIPS.
[63] Tuomas Haarnoja et al. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. ICML 2018.
[64] Kurtland Chua et al. (2018). Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models. NeurIPS.
[65] Alex Ray et al. (2019). Benchmarking Safe Exploration in Deep Reinforcement Learning. arXiv preprint.
[66] Tianhe Yu et al. (2020). Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning. CoRL.
[67] Shaoxiong Wang et al. (2022). Tacto: A Fast, Flexible, and Open-Source Simulator for High-Resolution Vision-Based Tactile Sensors. IEEE Robotics and Automation Letters.
[68] Zhenyu Wei et al. (2024). D(R,O) Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping. IEEE International Conference on Robotics and Automation (ICRA).
[69] Qian Mao et al. (2024). Multimodal Tactile Sensing Fused with Vision for Dexterous Robotic Housekeeping. Nature Communications.
[70] Chi, Cheng et al. (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv. #35 Terry commentary.
[71] Yang, Fengyu (2022). Touch and Go: Learning from Human-Collected Vision and Touch. arXiv.
[72] Li, Xingyu (2024). Evaluating Real-World Robot Manipulation Policies in Simulation. arXiv.
[73] Mittal, Mayank et al. (2025). Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning. arXiv.
[74] Black, Kevin (2024). pi0: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
[75] Octo Model Team (2024). Octo: An Open-Source Generalist Robot Policy. arXiv. #55 Terry commentary.
[76] Wu, Philipp et al. (2023). GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators. CoRL.
[77] Fu, Zipeng et al. (2024). Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation. arXiv.
[78] Cheng, Xuxin et al. (2024). Open-TeleVision: Teleoperation with Immersive Active Visual Feedback. CoRL.
[79] Zhaxizhuoma et al. (2024). FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset. arXiv.
[80] Wang, Chen et al. (2024). DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation. RSS.
[81] Xu, Chaoyi et al. (2026). RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning. arXiv. #36 Terry related commentary.
[82] Hoque, Ryan et al. (2025). EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. arXiv. #76 Terry commentary.
[83] Luo, Jianlan et al. (2025). Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning. Science Robotics.
[84] Xue, Zhengrong et al. (2025). DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning. arXiv.
[85] Jiang, Zhenyu et al. (2024). DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning. arXiv.
[86] Jang, Joel et al. (2025). DreamGen: Unlocking Generalization in Robot Learning through Video World Models. arXiv.
[87] Bousmalis, Konstantinos et al. (2023). RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation. arXiv.
[88] NVIDIA (2026a). NVIDIA Agent Toolkit Expands With New Omniverse Libraries, Putting AI Agents to Work Building Simulation-Ready Worlds. Official release.
[89] NVIDIA (2026b). Integrate Physical AI Capabilities into Existing Apps with NVIDIA Omniverse Libraries. NVIDIA Technical Blog.
[90] NVIDIA (2026c). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official technical video.
[91] NVIDIA (2026d). Bringing Agent-Ready Simulation Into Blender. Official technical video.
[92] Gao, Sicong et al. (2026). NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics. Third-party survey, arXiv.
[93] Fan, Jiacheng et al. (2026). RoboPaint: From Human Demonstration to Any Robot and Any View. arXiv. #15 Terry commentary.
[94] Zheng, Ruijie et al. (2026). EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data. arXiv.
[95] π0.6 Team (2025). π0.6: A Vision-Language-Action Model with Experience. Terry paper commentary [#4].
[96] ENPire Team (2026). ENPire: Robot Policy Self-Improvement. Terry paper commentary [#69].
[97] Jeffrey Mahler et al. (2017). Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics. Robotics: Science and Systems.
[98] Various (2024). DexH2R: Task-oriented Dexterous Manipulation from Human to Robots. arXiv preprint.
[99] Sungjae Park et al. (2025). Learning to Transfer Human Hand Skills for Robot Manipulations. arXiv preprint.
[100] Marc Peter Deisenroth et al. (2011). PILCO: A Model-Based and Data-Efficient Approach to Policy Search. ICML / Artificial Intelligence.
[101] Ankur Handa et al. (2020). DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System. ICRA 2020.
[102] Jeannette Bohg et al. (2014). Data-Driven Grasp Synthesis: A Survey. IEEE Transactions on Robotics.
[103] David J. Montana (1988). The Kinematics of Contact and Grasp. International Journal of Robotics Research.
[104] Genesis Team (2024). Genesis: A Generative and Universal Physics Engine for Robotics and Beyond. Project page.
[105] Hao, Peng et al. (2025). TLA: Tactile-Language-Action Model for Contact-Rich Manipulation. arXiv.
[106] Bjorck, Johan (2025). GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv.
[107] NVIDIA (2025). Isaac GR00T N1 Open Foundation Model for Humanoid Robots. NVIDIA Developer.
[108] DeepMind Robotics Team (2025). Gemini Robotics: Bringing AI into the Physical World. arXiv.
[109] Hogan, Neville (1985). Impedance Control: An Approach to Manipulation. Journal of Dynamic Systems, Measurement, and Control.
[110] Khatib, Oussama (1987). A Unified Approach for Motion and Force Control of Robot Manipulators: The Operational Space Formulation. IEEE Journal on Robotics and Automation.
[111] Mordatch, Igor et al. (2012). Discovery of Complex Behaviors through Contact-Invariant Optimization. ACM Transactions on Graphics.
[112] Posa, Michael et al. (2014). A Direct Method for Trajectory Optimization of Rigid Bodies Through Contact. The International Journal of Robotics Research.
[113] Todorov, Emanuel et al. (2012). MuJoCo: A Physics Engine for Model-Based Control. IEEE/RSJ IROS.
[114] Tobin, Josh et al. (2017). Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. IEEE/RSJ IROS.
[115] Peng, Xue Bin et al. (2018). Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. IEEE ICRA.
[116] Chebotar, Yevgen et al. (2019). Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience. IEEE ICRA.
[117] Siekmann, Jonah et al. (2021). Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning. RSS.
[118] He, Tairan et al. (2025). ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills. RSS.
[119] Ranawaka, Nadun et al. (2026). SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation. arXiv preprint.
[120] Ebert, Frederik et al. (2018). Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control. arXiv preprint.
[121] Gao, Shenyuan et al. (2026a). DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. ICML.
[122] Ye, Seonghyeon et al. (2026). World Action Models are Zero-shot Policies. arXiv preprint.
[123] NVIDIA Cosmos Team (2026). Cosmos 3: Omnimodal World Models for Physical AI. arXiv preprint.
[124] Gao, Sicong et al. (2026b). NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics. arXiv:2606.03551. Third-party survey.
[125] Nguyen, Duc Huy et al. (2024). TacEx: GelSight Tactile Simulation in Isaac Sim -- Combining Soft-Body and Visuotactile Simulators. arXiv preprint.
[126] NVIDIA (2026a). Newton Physics Engine. Official technical documentation.
[127] Alliance for OpenUSD (2026). OpenUSD Specifications and Alliance Governance. Official specification and governance site.
[128] NVIDIA (2026c). Bringing Agent-Ready Simulation Into Blender. Official demonstration video.
[129] NVIDIA (2026d). Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3. Official technical post.
[130] NVIDIA (2026e). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official tutorial video.
[131] Yuke Zhu et al. (2020). robosuite: A Modular Simulation Framework and Benchmark for Robot Learning. arXiv preprint.
[132] Rishabh Agarwal et al. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS.
[133] Ankur Handa et al. (2023). DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality. ICRA 2023.
[134] Zilin Si et al. (2024). DiffTactile: A Physics-based Differentiable Tactile Simulator for Contact-Rich Robotic Manipulation. ICLR 2024.
[135] Miquel Oller et al. (2024). Tactile-Driven Non-Prehensile Object Manipulation via Extrinsic Contact Mode Control. Robotics: Science and Systems (RSS) 2024.
[136] Various (2025). Robust Model-Based In-Hand Manipulation with Integrated Real-Time Motion-Contact Planning and Tracking. arXiv preprint.
[137] Uikyum Kim et al. (2021). Integrated Linkage-Driven Dexterous Anthropomorphic Robotic Hand. Nature Communications.
[138] Physical Intelligence (2025). OpenPI: Open Source Robot Policy Stack. GitHub.
[139] Ross, Stéphane et al. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
[140] Levine, Sergey et al. (2015). Learning Contact-Rich Manipulation Skills with Guided Policy Search. ICRA.
[141] Florence, Pete et al. (2021). Implicit Behavioral Cloning. CoRL.
[142] Kumar, Aviral et al. (2020). Conservative Q-Learning for Offline Reinforcement Learning. NeurIPS.
[143] Nair, Ashvin et al. (2020). AWAC: Accelerating Online Reinforcement Learning with Offline Datasets. arXiv.
[144] Yu, Tianhe et al. (2020). MOPO: Model-based Offline Policy Optimization. NeurIPS.
[145] Gulcehre, Caglar et al. (2020). RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning. NeurIPS.
[146] Luo, Jianlan et al. (2024). SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning. arXiv.
[147] Yu, Tianhe et al. (2021). MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale. arXiv.
[148] Lu, Yao et al. (2022). AW-Opt: Learning Robotic Skills with Imitation and Reinforcement at Scale. PMLR.
[149] Wu, Philipp et al. (2023). DayDreamer: World Models for Physical Robot Learning. CoRL.
[150] Achiam, Joshua et al. (2017). Constrained Policy Optimization. ICML.
[151] Rudin, Nikita et al. (2021). Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning. CoRL.
[154] Makoviychuk, Viktor et al. (2021). Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning. arXiv.
[155] Sergey Levine et al. (2016). End-to-End Training of Deep Visuomotor Policies. Journal of Machine Learning Research.
[156] Chelsea Finn et al. (2017). One-Shot Visual Imitation Learning via Meta-Learning. CoRL.
[157] John Schulman et al. (2017). Proximal Policy Optimization Algorithms. arXiv preprint.
[158] Andy Zeng et al. (2021). Transporter Networks: Rearranging the Visual World for Robotic Manipulation. CoRL.
[159] Eric Jang et al. (2022). BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning. CoRL.
[160] Niklas Funk et al. (2025). On the Importance of Tactile Sensing for Imitation Learning: A Case Study on Robotic Match Lighting. arXiv preprint.
[161] Anusha Nagabandi et al. (2018). Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning. IEEE ICRA.
[162] O'Neill, Abby et al. (2024). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. ICRA.
[163] Kim, Moo Jin (2024). OpenVLA: An Open-Source Vision-Language-Action Model. arXiv.
[164] Physical Intelligence (2024). pi0: A Generalist Robot Policy. Company research post.
[165] Pertsch, Karl (2025). FAST: Efficient Action Tokenization for Vision-Language-Action Models. arXiv.
[166] Shukor, Mustafa (2025). SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics. arXiv.
[167] Gemini Robotics Team (2025). Gemini Robotics: Bringing AI into the Physical World. arXiv.
[168] Yu, Jiawen (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv. #1 · KO
[169] Huang, Jialei (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv.
[170] Hao, Peng (2025). TLA: Tactile-Language-Action Model for Contact-Rich Manipulation. arXiv.
[171] AgiBot-World-Contributors (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
[172] Li, Xuanlin (2024). Evaluating Real-World Robot Manipulation Policies in Simulation. arXiv.
[173] Ross, Stephane et al. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
[174] Finn, Chelsea and Levine, Sergey (2017). Deep Visual Foresight for Planning Robot Motion. ICRA.
[175] Aditi et al. (2026). Cosmos 3: Omnimodal World Models for Physical AI. arXiv.
[176] Bhide, Asawaree et al. (2026). Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3. NVIDIA Technical Blog.
[177] NVIDIA Developer (2026). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official tutorial video.
[178] Li, Yang et al. (2026). ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation. arXiv.
[179] NVIDIA (2026). Cosmos3-Super Model Card. Hugging Face.
[180] Scott Fujimoto et al. (2018). Addressing Function Approximation Error in Actor-Critic Methods. ICML 2018.
[181] Jonathan Ho et al. (2020). Denoising Diffusion Probabilistic Models. NeurIPS 2020.
[182] Scott Reed et al. (2022). A Generalist Agent. Transactions on Machine Learning Research.
[183] Nur Muhammad Mahi Shafiullah et al. (2022). Behavior Transformers: Cloning k Modes with One Stone. NeurIPS.
[184] Mohit Shridhar et al. (2022). Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation. CoRL.
[185] Siddhant Haldar et al. (2024). BAKU: An Efficient Transformer for Multi-Task Policy Learning. NeurIPS 2024 / arXiv.
[186] Taowen Wang et al. (2024). Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics. Primary publication.
[187] Choi, Hojung et al. (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv. Terry #36 · KO.
[188] Xu, Mengda et al. (2025). DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation. arXiv. Terry #8 · KO.
[189] Fang, Hao-Shu et al. (2025). DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation. arXiv. Terry #10 · KO.
[190] Si, Zilin et al. (2025). ExoStart: Efficient learning for dexterous manipulation with sensorized exoskeleton demonstrations. arXiv. Terry #9 · KO.
[191] Qin, Yuzhe (2023). AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System. arXiv.
[192] Ding, Runyu (2024). Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning. arXiv.
[193] Kareer, Simar (2024). EgoMimic: Scaling Imitation Learning via Egocentric Video. arXiv.
[194] Yang, Fengyu (2023). Touch and Go: Learning from Human-Collected Vision and Touch. arXiv.
[195] AgiBot-World-Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
[196] Mandlekar, Ajay et al. (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
[197] Mandlekar, Ajay et al. (2018). RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation. arXiv.
[198] Zhao, Tony Z. et al. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
[199] Ohkawa, Takehiko et al. (2026). YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale. Technical report.
[200] Lum, Tyler Ga Wei et al. (2025). Crossing the Human-Robot Embodiment Gap with Sim-to-Real RL using One Human Demonstration. arXiv.
[201] Fu, Zipeng et al. (2024). HumanPlus: Humanoid Shadowing and Imitation from Humans. arXiv. Terry #38.
[202] Sundaram, Subramanian et al. (2019). Learning the Signatures of the Human Grasp Using a Scalable Tactile Glove. Nature.
[203] Yin, Jessica et al. (2025). OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer. arXiv. Terry #18 · KO.
[204] Dorsa Sadigh et al. (2017). Active Preference-Based Learning of Reward Functions. Robotics: Science and Systems.
[205] Roberto Calandra et al. (2018). More Than a Feeling: Learning to Grasp and Regrasp Using Vision and Touch. IEEE Robotics and Automation Letters.
[206] Runfa Blark Li et al. (2026). PhysGraph: Physically-Grounded Graph-Transformer Policies for Bimanual Dexterous Hand-Tool-Object Manipulation. arXiv preprint.
[207] Xingyu Peng et al. (2026). BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination. arXiv preprint.
[208] Alessio Palma et al. (2026). Bimanual Robot Manipulation via Multi-Agent In-Context Learning. arXiv preprint.
[209] Weiguang Zhao et al. (2026). Towards Robotic Dexterous Hand Intelligence: A Survey. arXiv preprint.
[210] OpenAI (2019). Learning Dexterous In-Hand Manipulation. International Journal of Robotics Research.
[211] Hogan, Neville (1985). Impedance Control: An Approach to Manipulation: Part I—Theory. Journal of Dynamic Systems, Measurement, and Control.
[212] Higuera, Carolina et al. (2024). Sparsh: Self-Supervised Touch Representations for Vision-Based Tactile Sensing. CoRL.
[213] Adeniji, Ademi et al. (2025). Feel the Force: Contact-Driven Learning from Humans. arXiv.
[214] Fengyu Yang et al. (2024). Binding Touch to Everything: Learning Unified Multimodal Tactile Representations. CVPR 2024.
[215] Sandra Q. Liu et al. (2024). A Passively Bendable, Compliant Tactile Palm with RObotic Modular Endoskeleton Optical (ROMEO) Fingers. IEEE International Conference on Robotics and Automation (ICRA).
[216] Jessica Yin et al. (2024). Learning In-Hand Translation Using Tactile Skin With Shear and Normal Force Sensing. arXiv preprint arXiv:2407.07885.
[217] Branden Romero et al. (2024). EyeSight Hand: Design of a Fully-Actuated Dexterous Robot Hand with Integrated Vision-Based Tactile Sensors and Compliant Actuation. IROS 2024.
[218] Binghao Huang et al. (2024). 3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing. CoRL 2024.
[219] Han Zhang et al. (2025). DOGlove: Dexterous Manipulation with a Low-Cost Open-Source Haptic Force Feedback Glove. RSS 2025.
[220] Akash Sharma et al. (2025). Self-supervised perception for tactile skin covered dexterous hands. arXiv preprint arXiv:2505.11420.
[221] Dasari, Sudeep et al. (2019). RoboNet: Large-Scale Multi-Robot Learning. CoRL.
[222] Open X-Embodiment Collaboration (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. IEEE ICRA 2024.
[223] Ghosh, Dibya et al. (2024). Octo: An Open-Source Generalist Robot Policy. RSS 2024.
[224] Kim, Moo Jin et al. (2024). OpenVLA: An Open-Source Vision-Language-Action Model. CoRL 2024.
[225] Black, Kevin et al. (2024). π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint. #2 · KO
[226] Bjorck, Johan et al. (2025). GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv preprint.
[227] Physical Intelligence et al. (2025). π0.5: A Vision-Language-Action Model with Open-World Generalization. arXiv preprint.
[228] Physical Intelligence (2026). $π_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv preprint. #62 · KO
[229] NVIDIA Omniverse (2026). NVIDIA Agent Toolkit Expands With New Omniverse Libraries. Official announcement.
[230] Newton Project (2026). Newton Physics Engine. Official project page.
[231] NVIDIA Cosmos (2026). Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3. Official technical blog.
[232] NVIDIA Video (2026a). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official tutorial video.
[233] NVIDIA Video (2026b). Bringing Agent-Ready Simulation Into Blender. Official workflow video.
[234] Yu, Jiawen et al. (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv preprint.
[235] Huang, Jialei et al. (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv preprint.
[237] Figure AI (2025). Helix: A Vision-Language-Action Model for Generalist Humanoid Control. Official company research post.
[238] Generalist Team (2026). GEN-1: Scaling Embodied Foundation Models to Mastery. Official company research post. #24 · KO
[239] Chelsea Finn et al. (2016). Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization. International Conference on Machine Learning.
[240] Michael Janner et al. (2019). When to Trust Your Model: Model-Based Policy Optimization. NeurIPS.
[241] Oliver Kroemer et al. (2021). A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms. JMLR.
[242] Fanbo Xiang et al. (2020). SAPIEN: A SimulAted Part-based Interactive ENvironment. CVPR.
[243] Mohit Shridhar et al. (2022). CLIPort: What and Where Pathways for Robotic Manipulation. CoRL.
[244] Songming Liu et al. (2024). RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation. arXiv primary preprint (2024-10-10).
[245] Pertsch, Karl et al. (2024). Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models. Primary publication.
[246] Team AgiBot-World (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
[247] Ross, Stéphane, Gordon, Geoffrey J., and Bagnell, J. Andrew (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
[248] Kalashnikov, Dmitry et al. (2021). MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale. arXiv.
[249] Dexterity (2026). Introducing Instinct. Company engineering report.
[250] Aravind Rajeswaran et al. (2018). Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. Robotics: Science and Systems.
[251] Tongzhou Mu et al. (2021). ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations. NeurIPS Datasets and Benchmarks.
[252] Suraj Nair et al. (2022). R3M: A Universal Visual Representation for Robot Manipulation. CoRL.
[253] Jiayuan Gu et al. (2023). ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills. ICLR 2023.
[254] Yongchao Chen et al. (2025). Code-as-Symbolic-Planner: Foundation Model-Based Robot Planning via Symbolic Code Generation. IROS.
[255] Zhongxuan Li et al. (2026). UniBiDex: A Unified Teleoperation Framework for Robotic Bimanual Dexterous Manipulation. arXiv preprint.
[256] Javier Romero et al. (2017). MANO: A Hand Model with Articulated and Non-rigid Deformations. SIGGRAPH Asia 2017.
[257] Si, Zilin (2025). ExoStart: From 10 Exoskeleton Demos to Dexterous Robot Manipulation. arXiv.
[258] Zhao, C. et al. (2025a). Universal Slip Detection of Robotic Hand with Tactile Sensing. Frontiers in Neurorobotics.
[259] Hao, Yaru (2025). Tactile-Language-Action Model for Contact-Rich Manipulation. arXiv.
[260] Generalist AI (2025). GEN-0 Robot Foundation Model. Company page.
[261] Skild AI (2024). General-Purpose Robot Brain. Company page.
[262] Zhao, Zihang et al. (2025b). Embedding High-Resolution Touch across Robotic Hands Enables Adaptive Human-like Grasping. Nature Machine Intelligence.
[263] Zhang, Ningbin et al. (2025a). Soft Robotic Hand with Tactile Palm-Finger Coordination. Nature Communications.
[264] Catalano, Manuel G. et al. (2014). Adaptive Synergies for the Design and Control of the Pisa/IIT SoftHand. The International Journal of Robotics Research.
[265] Della Santina, Cosimo et al. (2018). Toward Dexterous Manipulation with Augmented Adaptive Synergies: The Pisa/IIT SoftHand 2. IEEE Transactions on Robotics.
[266] Zhang, Chi et al. (2025b). UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer. arXiv preprint. #16
[267] Fang, Hongjie et al. (2024). AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild. arXiv.
[268] Bhirangi, Raunaq et al. (2021). ReSkin: Versatile, Replaceable, Lasting Tactile Skins. arXiv.
[269] Christoph, Clemens C. et al. (2025). ORCA: An Open-Source, Reliable, Cost-Effective, Anthropomorphic Robotic Hand for Uninterrupted Dexterous Task Learning. arXiv.
[270] Open X-Embodiment Collaboration et al. (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
[271] Chi, Cheng (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv.
[272] Covariant (2024). Covariant Introduces RFM-1 to Give Robots the Human-like Ability to Reason. First-party-distributed company release.
[273] NVIDIA (2026). NVIDIA Agent Toolkit Expands With New Omniverse Libraries. First-party release.
[274] Toyota Research Institute (2023). Toyota Research Institute Unveils Breakthrough in Teaching Robots New Behaviors. Company technical post.
[275] Dulac-Arnold, Gabriel et al. (2021). Challenges of Real-World Reinforcement Learning: Definitions, Benchmarks and Analysis. Machine Learning.
[276] Dan, Prithwish et al. (2025). X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real. arXiv.
[277] Xiao, Wenli et al. (2026). ENPIRE: Agentic Robot Policy Self-Improvement in the Real World. arXiv. #69
Acknowledgment
This survey connects prior work from S1 robot hands, S4 humanoids, S6 manufacturing physical AI, and S9 NVIDIA physical AI, while treating large-data driven manipulation as its own subject.
This project was built using the Harness skill by Minho Hwang.
AI tools were used in the production of this work: Claude (Opus 4.6) for literature survey, content generation, and manuscript preparation.