통합 참고문헌 (References)

277 references

[2] Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
[4] Dasari, Sudeep (2019). RoboNet: Large-Scale Multi-Robot Learning. arXiv.
[11] Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv.
[16] Covariant (2024). RFM-1: Robotics Foundation Model. Company technical post.
[17] Dexterity (2025). Dexterity Foresight: AI Platform for Industrial Robot Workcells. Company product page.
[18] Chef Robotics (2025). ChefOS: AI Robotics Platform for Food Manufacturing. Company product page.
[19] Toyota Research Institute (2024). Large Behavior Models for Robot Manipulation. Company technical post.
[28] Argall, Brenna D. et al. (2009). A Survey of Robot Learning from Demonstration. Robotics and Autonomous Systems.
[31] Billard, Aude and Kragic, Danica (2019). Trends and Challenges in Robot Manipulation. Science.
[35] Corey Lynch et al. (2020). Learning Latent Plans from Play. CoRL.
[36] Tony Z. Zhao et al. (2024). ALOHA Unleashed: A Simple Recipe for Robot Dexterity. Conference on Robot Learning (CoRL).
[37] Kailin Li et al. (2025). ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025.
[38] Roland S. Johansson et al. (2009). Coding and Use of Tactile Signals from the Fingertips in Object Manipulation Tasks. Nature Reviews Neuroscience.
[39] Chelsea Finn et al. (2017). Deep Visual Foresight for Planning Robot Motion. IEEE ICRA.
[40] Luis Sentis et al. (2005). Synthesis of Whole-Body Behaviors through Hierarchical Control of Behavioral Primitives. International Journal of Humanoid Robotics.
[41] Matthew T. Mason (1986). Mechanics and Planning of Manipulator Pushing Operations. International Journal of Robotics Research.
[44] Bhirangi, Raunaq (2021). ReSkin: versatile, replaceable, lasting tactile skins. arXiv.
[45] Bhirangi, Raunaq (2024). AnySkin: Plug-and-play Skin Sensing for Robotic Touch. arXiv.
[52] Figure AI (2026). Figure 03 + Helix 02: General-Purpose Humanoid System. Company product page.
[53] Physical Intelligence (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. Company research post.
[54] Bicchi, Antonio (2000). Hands for Dexterous Manipulation and Robust Grasping: A Difficult Road Toward Simplicity. IEEE Transactions on Robotics and Automation.
[55] Capsi-Morales, Patricia et al. (2020). Exploring the Role of Palm Concavity and Adaptability in Soft Synergistic Robotic Hands. IEEE Robotics and Automation Letters.
[56] Agarwal, Arpit et al. (2025). A Modularized Design Approach for GelSight Family of Vision-based Tactile Sensors. The International Journal of Robotics Research.
[57] Fang, Bin et al. (2025). Force Measurement Technology of Vision-Based Tactile Sensor. Advanced Intelligent Systems.
[58] Gupta, Harsh et al. (2025). Sensor-Invariant Tactile Representation. arXiv.
[59] Zhao, Zihang et al. (2025). Embedding High-Resolution Touch across Robotic Hands Enables Adaptive Human-Like Grasping. Nature Machine Intelligence.
[62] Felix Berkenkamp et al. (2017). Safe Model-based Reinforcement Learning with Stability Guarantees. NeurIPS.
[65] Alex Ray et al. (2019). Benchmarking Safe Exploration in Deep Reinforcement Learning. arXiv preprint.
[67] Shaoxiong Wang et al. (2022). Tacto: A Fast, Flexible, and Open-Source Simulator for High-Resolution Vision-Based Tactile Sensors. IEEE Robotics and Automation Letters.
[68] Zhenyu Wei et al. (2024). D(R,O) Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping. IEEE International Conference on Robotics and Automation (ICRA).
[69] Qian Mao et al. (2024). Multimodal Tactile Sensing Fused with Vision for Dexterous Robotic Housekeeping. Nature Communications.
[87] Bousmalis, Konstantinos et al. (2023). RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation. arXiv.
[90] NVIDIA (2026c). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. 공식 기술 영상.
[91] NVIDIA (2026d). Bringing Agent-Ready Simulation Into Blender. 공식 기술 영상.
[92] Gao, Sicong et al. (2026). NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics. 제3자 survey, arXiv.
[95] π0.6 Team (2025). π0.6: A Vision-Language-Action Model with Experience. Terry의 논문 해설 [#4].
[96] ENPire Team (2026). ENPire: Robot Policy Self-Improvement. Terry의 논문 해설 [#69].
[99] Sungjae Park et al. (2025). Learning to Transfer Human Hand Skills for Robot Manipulations. arXiv preprint.
[100] Marc Peter Deisenroth et al. (2011). PILCO: A Model-Based and Data-Efficient Approach to Policy Search. ICML / Artificial Intelligence.
[102] Jeannette Bohg et al. (2014). Data-Driven Grasp Synthesis: A Survey. IEEE Transactions on Robotics.
[103] David J. Montana (1988). The Kinematics of Contact and Grasp. International Journal of Robotics Research.
[107] NVIDIA (2025). Isaac GR00T N1 Open Foundation Model for Humanoid Robots. NVIDIA Developer.
[108] DeepMind Robotics Team (2025). Gemini Robotics: Bringing AI into the Physical World. arXiv.
[109] Hogan, Neville (1985). Impedance Control: An Approach to Manipulation. Journal of Dynamic Systems, Measurement, and Control.
[110] Khatib, Oussama (1987). A Unified Approach for Motion and Force Control of Robot Manipulators: The Operational Space Formulation. IEEE Journal on Robotics and Automation.
[111] Mordatch, Igor et al. (2012). Discovery of Complex Behaviors through Contact-Invariant Optimization. ACM Transactions on Graphics.
[112] Posa, Michael et al. (2014). A Direct Method for Trajectory Optimization of Rigid Bodies Through Contact. The International Journal of Robotics Research.
[113] Todorov, Emanuel et al. (2012). MuJoCo: A Physics Engine for Model-Based Control. IEEE/RSJ IROS.
[115] Peng, Xue Bin et al. (2018). Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. IEEE ICRA.
[122] Ye, Seonghyeon et al. (2026). World Action Models are Zero-shot Policies. arXiv preprint.
[123] NVIDIA Cosmos Team (2026). Cosmos 3: Omnimodal World Models for Physical AI. arXiv preprint.
[124] Gao, Sicong et al. (2026b). NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics. arXiv:2606.03551. Third-party survey.
[126] NVIDIA (2026a). Newton Physics Engine. Official technical documentation.
[127] Alliance for OpenUSD (2026). OpenUSD Specifications and Alliance Governance. Official specification and governance site.
[128] NVIDIA (2026c). Bringing Agent-Ready Simulation Into Blender. Official demonstration video.
[130] NVIDIA (2026e). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official tutorial video.
[132] Rishabh Agarwal et al. (2021). Deep Reinforcement Learning at the Edge of the Statistical Precipice. NeurIPS.
[135] Miquel Oller et al. (2024). Tactile-Driven Non-Prehensile Object Manipulation via Extrinsic Contact Mode Control. Robotics: Science and Systems (RSS) 2024.
[137] Uikyum Kim et al. (2021). Integrated Linkage-Driven Dexterous Anthropomorphic Robotic Hand. Nature Communications.
[138] Physical Intelligence (2025). OpenPI: Open Source Robot Policy Stack. GitHub.
[141] Florence, Pete et al. (2021). Implicit Behavioral Cloning. CoRL.
[142] Kumar, Aviral et al. (2020). Conservative Q-Learning for Offline Reinforcement Learning. NeurIPS.
[144] Yu, Tianhe et al. (2020). MOPO: Model-based Offline Policy Optimization. NeurIPS.
[145] Gulcehre, Caglar et al. (2020). RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning. NeurIPS.
[149] Wu, Philipp et al. (2023). DayDreamer: World Models for Physical Robot Learning. CoRL.
[150] Achiam, Joshua et al. (2017). Constrained Policy Optimization. ICML.
[152] Physical Intelligence (2025). π*₀.₆: a VLA That Learns From Experience. arXiv. #4
[153] Terry Um (2026). ENPIRE: Robot Policy Self-Improvement. Terry's Blog. #69
[155] Sergey Levine et al. (2016). End-to-End Training of Deep Visuomotor Policies. Journal of Machine Learning Research.
[156] Chelsea Finn et al. (2017). One-Shot Visual Imitation Learning via Meta-Learning. CoRL.
[157] John Schulman et al. (2017). Proximal Policy Optimization Algorithms. arXiv preprint.
[162] O'Neill, Abby et al. (2024). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. ICRA.
[163] Kim, Moo Jin (2024). OpenVLA: An Open-Source Vision-Language-Action Model. arXiv.
[164] Physical Intelligence (2024). pi0: A Generalist Robot Policy. Company research post.
[167] Gemini Robotics Team (2025). Gemini Robotics: Bringing AI into the Physical World. arXiv.
[174] Finn, Chelsea and Levine, Sergey (2017). Deep Visual Foresight for Planning Robot Motion. ICRA.
[175] Aditi et al. (2026). Cosmos 3: Omnimodal World Models for Physical AI. arXiv.
[176] Bhide, Asawaree et al. (2026). Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3. NVIDIA Technical Blog.
[177] NVIDIA Developer (2026). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official tutorial video.
[179] NVIDIA (2026). Cosmos3-Super Model Card. Hugging Face.
[180] Scott Fujimoto et al. (2018). Addressing Function Approximation Error in Actor-Critic Methods. ICML 2018.
[181] Jonathan Ho et al. (2020). Denoising Diffusion Probabilistic Models. NeurIPS 2020.
[182] Scott Reed et al. (2022). A Generalist Agent. Transactions on Machine Learning Research.
[183] Nur Muhammad Mahi Shafiullah et al. (2022). Behavior Transformers: Cloning k Modes with One Stone. NeurIPS.
[185] Siddhant Haldar et al. (2024). BAKU: An Efficient Transformer for Multi-Task Policy Learning. NeurIPS 2024 / arXiv.
[187] Choi, Hojung et al. (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv. Terry #36.
[202] Sundaram, Subramanian et al. (2019). Learning the Signatures of the Human Grasp Using a Scalable Tactile Glove. Nature.
[204] Dorsa Sadigh et al. (2017). Active Preference-Based Learning of Reward Functions. Robotics: Science and Systems.
[205] Roberto Calandra et al. (2018). More Than a Feeling: Learning to Grasp and Regrasp Using Vision and Touch. IEEE Robotics and Automation Letters.
[208] Alessio Palma et al. (2026). Bimanual Robot Manipulation via Multi-Agent In-Context Learning. arXiv preprint.
[209] Weiguang Zhao et al. (2026). Towards Robotic Dexterous Hand Intelligence: A Survey. arXiv preprint.
[210] OpenAI (2019). Learning Dexterous In-Hand Manipulation. International Journal of Robotics Research.
[211] Hogan, Neville (1985). Impedance Control: An Approach to Manipulation: Part I—Theory. Journal of Dynamic Systems, Measurement, and Control.
[213] Adeniji, Ademi et al. (2025). Feel the Force: Contact-Driven Learning from Humans. arXiv.
[215] Sandra Q. Liu et al. (2024). A Passively Bendable, Compliant Tactile Palm with RObotic Modular Endoskeleton Optical (ROMEO) Fingers. IEEE International Conference on Robotics and Automation (ICRA).
[216] Jessica Yin et al. (2024). Learning In-Hand Translation Using Tactile Skin With Shear and Normal Force Sensing. arXiv preprint arXiv:2407.07885.
[220] Akash Sharma et al. (2025). Self-supervised perception for tactile skin covered dexterous hands. arXiv preprint arXiv:2505.11420.
[221] Dasari, Sudeep et al. (2019). RoboNet: Large-Scale Multi-Robot Learning. CoRL.
[222] Open X-Embodiment Collaboration (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. IEEE ICRA 2024.
[223] Ghosh, Dibya et al. (2024). Octo: An Open-Source Generalist Robot Policy. RSS 2024.
[224] Kim, Moo Jin et al. (2024). OpenVLA: An Open-Source Vision-Language-Action Model. CoRL 2024.
[225] Black, Kevin et al. (2024). π0: A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint. #2
[226] Bjorck, Johan et al. (2025). GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv preprint.
[227] Physical Intelligence et al. (2025). π0.5: A Vision-Language-Action Model with Open-World Generalization. arXiv preprint.
[229] NVIDIA Omniverse (2026). NVIDIA Agent Toolkit Expands With New Omniverse Libraries. Official announcement.
[230] Newton Project (2026). Newton Physics Engine. Official project page.
[231] NVIDIA Cosmos (2026). Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3. Official technical blog.
[232] NVIDIA Video (2026a). How to Post-Train NVIDIA Cosmos 3 for Robot Action Prediction. Official tutorial video.
[233] NVIDIA Video (2026b). Bringing Agent-Ready Simulation Into Blender. Official workflow video.
[236] Kim, Dongyoung et al. (2026). RLDX-1 Technical Report. arXiv preprint. #67
[237] Figure AI (2025). Helix: A Vision-Language-Action Model for Generalist Humanoid Control. Official company research post.
[238] Generalist Team (2026). GEN-1: Scaling Embodied Foundation Models to Mastery. Official company research post. #24
[239] Chelsea Finn et al. (2016). Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization. International Conference on Machine Learning.
[240] Michael Janner et al. (2019). When to Trust Your Model: Model-Based Policy Optimization. NeurIPS.
[242] Fanbo Xiang et al. (2020). SAPIEN: A SimulAted Part-based Interactive ENvironment. CVPR.
[243] Mohit Shridhar et al. (2022). CLIPort: What and Where Pathways for Robotic Manipulation. CoRL.
[244] Songming Liu et al. (2024). RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation. arXiv primary preprint (2024-10-10).
[247] Ross, Stéphane, Gordon, Geoffrey J., and Bagnell, J. Andrew (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
[248] Kalashnikov, Dmitry et al. (2021). MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale. arXiv.
[249] Dexterity (2026). Introducing Instinct. Company engineering report.
[250] Aravind Rajeswaran et al. (2018). Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. Robotics: Science and Systems.
[251] Tongzhou Mu et al. (2021). ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations. NeurIPS Datasets and Benchmarks.
[256] Javier Romero et al. (2017). MANO: A Hand Model with Articulated and Non-rigid Deformations. SIGGRAPH Asia 2017.
[258] Zhao, C. et al. (2025a). Universal Slip Detection of Robotic Hand with Tactile Sensing. Frontiers in Neurorobotics.
[260] Generalist AI (2025). GEN-0 Robot Foundation Model. Company page.
[261] Skild AI (2024). General-Purpose Robot Brain. Company page.
[262] Zhao, Zihang et al. (2025b). Embedding High-Resolution Touch across Robotic Hands Enables Adaptive Human-like Grasping. Nature Machine Intelligence.
[263] Zhang, Ningbin et al. (2025a). Soft Robotic Hand with Tactile Palm-Finger Coordination. Nature Communications.
[264] Catalano, Manuel G. et al. (2014). Adaptive Synergies for the Design and Control of the Pisa/IIT SoftHand. The International Journal of Robotics Research.
[265] Della Santina, Cosimo et al. (2018). Toward Dexterous Manipulation with Augmented Adaptive Synergies: The Pisa/IIT SoftHand 2. IEEE Transactions on Robotics.
[268] Bhirangi, Raunaq et al. (2021). ReSkin: Versatile, Replaceable, Lasting Tactile Skins. arXiv.
[270] Open X-Embodiment Collaboration et al. (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
[272] Covariant (2024). Covariant Introduces RFM-1 to Give Robots the Human-like Ability to Reason. First-party-distributed company release.
[273] NVIDIA (2026). NVIDIA Agent Toolkit Expands With New Omniverse Libraries. First-party release.
[274] Toyota Research Institute (2023). Toyota Research Institute Unveils Breakthrough in Teaching Robots New Behaviors. Company technical post.
[275] Dulac-Arnold, Gabriel et al. (2021). Challenges of Real-World Reinforcement Learning: Definitions, Benchmarks and Analysis. Machine Learning.
[276] Dan, Prithwish et al. (2025). X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real. arXiv.

감사의 글

이 서베이는 S1 로봇 핸드, S4 휴머노이드, S6 제조 피지컬AI, S9 NVIDIA 피지컬AI의 논지를 연결하되, large-data driven manipulation 자체를 독립 주제로 재구성한다.

이 프로젝트는 황민호님의 Harness 스킬을 이용하여 제작되었습니다.

이 저작물의 제작에 AI 도구가 활용되었습니다. 문헌 조사, 콘텐츠 생성, 원고 작성에 Claude(Opus 4.6)를 사용하였습니다.