Chapter 2: End Effectors — Designing Data Difficulty
Overview
The end effector is not merely the last component of a robot. It is the coordinate system where data meets the world. Suction, two-finger grippers, custom tools, and five-finger hands define different executable actions, expose different failures, and create different retraining costs. Changing the hand therefore changes command fields, sensor streams, failure labels, replay tests, and maintenance history. LEAP Hand and DIGIT-style tactile fingertips broaden access to low-cost hands and touch sensors [1] [2]. In a manufacturing cell, however, resemblance to the human hand matters less than repeatable evidence about failure.
This chapter chooses the hand through four questions. How narrowly should the task be partitioned? How directly should contact be measured? How will human demonstrations transfer into the target embodiment? Can sensor and hand maintenance fit inside production cost? Work on ReSkin, AnySkin, CoinFT, AnyTouch, and ForceVLA shows that surface sensing, force, and tactile representation are becoming part of the policy interface [4] [5] [6] [7] [13]. Their results remain bounded to particular hands, tasks, sensors, data splits, and controllers. This chapter preserves those boundaries and adds the manufacturing interpretation.
Learning objectives - Explain end-effector choice as the joint design of an action space and an observation space. - Compare suction, two fingers, custom tools, and five-finger hands by failure observability, maintenance burden, and reproducibility. - Distinguish conditions in which touch or force reduces task ambiguity from conditions in which it merely adds calibration burden. - Explain why human demonstrations and target robot hands can assign different meanings to the same contact. - Specify why an end-effector change also changes the episode schema, replay suite, and maintenance history.
The Hand Is the First Task Classifier
It is risky to assume that a large policy will treat all hands as interchangeable. The end-effector defines the action space and the failures that can be observed. Suction leaves strong signals about seal and vacuum pressure, but weak evidence about fine part rotation or edge damage. A two-finger gripper records width, closing force, and regrasp more cleanly, but may miss distributed contact on deformable objects. A five-finger hand creates richer contact, but calibration, cleaning, and repair logs become part of the episode.
Figure 2.1. End-effector taxonomy as data strategy. Hand choice partitions the task family, sensor channels, and failure labels at the same time. Source: author-created SVG.
| End-effector | Good factory fit | Strong data signal | Main blind spot |
|---|---|---|---|
| Suction | Flat packaging, trays, cartons, fast picking | Vacuum pressure, seal break, pick/no-pick | Contact location, fine rotation, deformable damage |
| Two-finger gripper | Small parts, caps, simple insertion, regrasp | Width, force/torque, jaw pose, retry | Surface slip, distributed contact |
| Custom hand/tool | Screw, wipe, press, fixture-specific work | Process-specific force and jig state | Reuse when the task changes |
| Five-finger hand | Tool use, cable, fabric, irregular objects | Multi-contact, in-hand motion, human-like demos | Maintenance, sensor drift, data dimensionality |
The table is not meant to crown a universal best hand. The manufacturer should first identify which failure costs the most. If scratches dominate, tactile skin and force limits may matter most. If missed picks dominate, suction pressure and camera coverage may be more valuable.
End-effector mechanics determine both the executable actions and the physical signals available for diagnosing failure. That statement does not mean fewer degrees of freedom are universally better. Bicchi [25] summarized an early lesson from multi-finger hands: robustness and mechanical simplicity can matter more than degree-of-freedom count. Capsi-Morales et al.[2] later showed, on a particular SoftHand and object set, that palm concavity and compliance change the envelope of successful grasping. These results support the thesis that mechanics shapes data difficulty. They do not establish universal transfer across hands, independent production readiness, or lower total data cost. Those outcomes must be tested outside the cited hand, object set, controller, and evaluation denominator.
Historical Transition: From Degree-of-Freedom Competition to Co-Design
The history of robot hands is not a monotonic march toward more joints. Early anthropomorphic hands reproduced more of the human hand's kinematics, but each added actuator and transmission path also added control variables, contact error, and service points. Bicchi [25] framed this difficult road as a search for simplicity. Compliant, underactuated, and synergy-driven designs subsequently showed how a hand could adapt to object shape without commanding every joint independently. In data terms, mechanical structure became an inductive bias that removed variation a policy would otherwise have to learn.
Learning-based robotics did not invalidate that principle. It made the tradeoff easier to see. More joint commands enlarge the action space covered by a fixed number of episodes, while additional contacts multiply force and location combinations. Yet a hand can also be too simple to express the state transitions required by cable twisting, cloth folding, tool regrasping, or in-hand rotation. The useful question is therefore not “How many degrees of freedom does the hand have?” It is “What is the smallest action coordinate system that expresses success, and what is the smallest observation coordinate system that explains failure?”
This perspective does not put suction and a five-finger hand on one universal ladder. Suction creates a compact data contract around seal state, vacuum pressure, and release timing. A five-finger hand creates a wider contract involving per-finger position, distributed contact, slip, tendon state, and recovery. Their value depends on defect cost, product variety, replacement time, and operator intervention. Saying that mechanics is part of the policy is not an instruction to make hardware more complex. It is an instruction to allocate necessary complexity deliberately between mechanics and learning.
Why Two Fingers and Suction Still Win Often
In much factory automation, repeatability beats dexterity. Suction and two-finger grippers are simple, which makes episodes shorter, fault trees smaller, and maintenance more predictable. They also help large-data learning. A smaller action space gives faster coverage for the same number of episodes and makes it easier to separate hardware faults from policy faults.
Simplicity is not free. Suction provides a strong picked/not-picked signal but weak evidence about how an object deformed during contact. A two-finger gripper records force and pose well, but soft packaging or cloth can create wide contact patches that the jaws do not explain. To decide whether suction or two fingers are enough, inspect the reject codes instead of the success clips. The hand is chosen by the most expensive failure: missed pick, crush, scratch, wrong orientation, or downstream jam.
Tactile Fingertips Are Part of the Hand
A tactile sensor is not a decorative add-on; it changes the observation surface of the hand. DIGIT, ReSkin, AnySkin, and Wedge-style sensors show a trend toward touch sensing that is cheaper, replaceable, and more geometry-aware [2] [4] [5]. This matters in manufacturing because a fragile sensor can create more downtime than data value.
Figure 2.2. GelSight Wedge shows tactile geometry for robot fingers. Tactile geometry turns contact patches into data, but cleaning and replacement cycles must be designed with it. Source: Wang et al. 2021, arXiv:2106.08851 Fig. 1.
Touch is most useful when vision sees the failure too late. The moment an insertion binds, a cable begins to slip, or a pouch surface folds is often visible first in contact change rather than in camera frames. AnyTouch and Tactile-VLA point toward tactile representations that can transfer across sensors and policies [7] [14] [15]. A manufacturer still has to place learning value and operating cost in the same decision table.
High-resolution tactile sensors can expose contact geometry and force cues that are not directly available to a vision-only policy. Vision-based tactile sensors turn deformation of an elastomer into images that can encode contact patches, shear, and slip proxies. Modular GelSight work makes optics, illumination, gel, and housing part of a task-specific design space [27]. Work on force measurement explains why optical geometry, elastomer mechanics, and calibration models are all required to infer force from those images [28]. A camera inside a fingertip does not reveal contact ground truth by itself. Change the geometry, illumination, gel wear, or estimator training distribution and the meaning of its output can change.
Observability Means Failure Resolution, Not Channel Count
Counting cameras and force axes makes instrumentation easy to purchase, but manufacturing needs enough resolution to distinguish causes. Suppose an insertion failure can result from pose error, a defective chamfer, increased friction, or an inverted part. The team must ask whether wrist force separates those four causes. If it does not, collecting more episodes may attach different labels to the same observation. That is not primarily a learning problem. It is an identifiability problem in the measurement design.
Vision and touch are not competing modalities. An external camera can observe the approach path and global part pose. A wrist force-torque sensor observes aggregate contact load. A tactile fingertip can observe a local patch or the onset of slip. A useful design assigns each branch in the defect tree to the lowest-cost signal that can separate it: vision before contact, force at first contact, and touch during fine alignment, for example. For repetitive lifting of flat cartons, however, vacuum pressure and one camera may explain failure better than a dexterous hand with many channels.
Sensor usefulness depends on calibration, synchronization, durability, cleaning, replacement, and recalibration rather than nominal channel count alone. ReSkin established replaceability as an explicit tactile-skin objective [4], while sensor-invariant representation learning explores common features across tactile devices [29]. Neither direction erases maintenance history. The cited sensor and task results do not independently establish long-duration drift behavior, equivalence after repeated replacement, contamination tolerance, or production uptime. If zero, sensitivity, delay, and surface state are not recorded before and after replacement, rich tactile data becomes a rich confounder.
Connecting the sensor lifecycle to episodes requires at least a sensor identifier, material batch, calibration time, residual error, sample rate, clock offset, cleaning procedure, and cumulative contact count. If storing every raw stream indefinitely is impractical, the cell can retain event windows and summary statistics around failures. The critical requirement is that policy input and quality investigation refer to the same sensor state. When the sensor changes but the dataset version does not, no analysis can reliably separate algorithm improvement from instrumentation change.
Decision Walkthrough: Tray Picking and Cable Routing Need Different Hands
Consider a packaging-tray pick cell. The objects are relatively flat, and the failures are missed pick, double pick, crush, and contamination. Suction is a strong baseline. Vacuum pressure, cup wear, release timing, and camera pose explain many of the expensive failures. A five-finger hand could create richer contact, but it also introduces cleaning, food-safety, residue, and calibration risks. A more complex hand is not automatically a better data instrument.
Cable routing is different. The cable changes shape while moving, slips between fingers, catches on fixture edges, and produces abrupt force changes near final insertion. Suction explains almost none of that state. A simple two-finger gripper may also be too weak if jaw force is the only contact signal. The cell may need tactile fingertips, a compliant fixture, or a multi-contact hand. The key question is not whether the hand is human-like, but whether it preserves the state transitions that explain failure.
| Task example | First hand to test | Additional sensor candidate | Hand-upgrade trigger |
|---|---|---|---|
| Food-tray picking | Suction or soft two-finger gripper | Vacuum, vision, cup-wear log | Crush or contamination becomes more expensive than missed pick |
| Cap handling | Two-finger gripper | Wrist force/torque, jaw width | Cross-thread and slip cannot be separated |
| Cable routing | Two fingers with tactile or multi-contact hand | Tactile patch, cable tension | Jaw force cannot localize the jam |
| Fabric or pouch manipulation | Custom fixture plus tactile hand | Distributed touch, deformation cue | Surface damage and fold failures repeat |
This walkthrough turns hardware selection into task taxonomy. A tray cell and a cable-routing cell can both be part of a large-data strategy, but the former is governed by throughput and cleaning evidence, while the latter is governed by contact observability. The manufacturer should ask which expensive failure disappears under each hand, not whether a five-finger hand is available.
When Five Fingers Become Justified
A five-finger hand is not justified because it looks human. It is justified when the work value requires extra contact freedom and that freedom reduces actual failure labels. Cable routing, tool handoff, deformable packaging, in-hand reorientation, and human tool-use transfer can qualify. LEAP Hand, DexUMI, DEXOP, and ExoStart show how low-cost dexterous hands and human-hand transfer are advancing [1] [8] [28] [10].
The cost is a data explosion. Joint state, fingertip contact, tendon and friction state, calibration, wear, and recovery behavior all enter the episode. As the hand grows more complex, the operating schema becomes more important than the policy architecture. Can the same model be deployed after a hand repair? Do glove or exoskeleton demonstrations preserve the same contact semantics as the robot hand? How should a replay set mark one drifting fingertip sensor? Without answers, dexterity becomes debugging cost.
Distributed tactile coverage can improve adaptive grasping within a reported hand and task setup, but it does not by itself prove transfer across hands. Zhao et al.[6] reported adaptive grasping with high-resolution touch distributed across a hand, and Almeida et al.[6] studied how tactile placement relates to in-hand manipulation. Zhang et al.[6] explore cross-hand tactile representation in UniTacHand, but the evidence is bounded by a small paired dataset and one robot platform. The manufacturing conclusion is not “more tactile cells guarantee generalization.” It is “measure which contact regions identify the state required by this task.”
The frontier contribution is therefore not that five fingers have become the default. It is the recognition that sensor placement, hand mechanics, and representation learning cannot be optimized independently. A policy cannot recover an unobserved state if only fingertips are instrumented when the task depends on palm contact, or if only wrist load is recorded when fingertip slip is decisive. A cell can begin with sensor ablation: mask one channel at a time and measure the change in defect classification and recovery. Channels with little marginal value become removal candidates once maintenance burden is included.
Force/Torque Makes Simple Hands More Honest
Not every cell needs five fingers. A two-finger gripper or compact custom tool with force/torque sensing can be the better manufacturing choice. CoinFT-style compact six-axis force-torque sensing shows how a small sensor can make a simple gripper produce contact-rich episodes [6].
Figure 2.3. Compact force-torque sensing changes the hand data contract. Small force/torque sensors improve failure explanation for simple grippers and custom tools. Source: Choi et al. 2025, arXiv:2503.19225 Fig. 1.
In cap fastening, for example, a two-finger gripper with a torque-aware wrist may be better than a five-finger hand. The task needs signals for thread start, cross-thread, over-torque, and cap slip, not a human-shaped grasp. Cable routing is different: jaw force alone may not explain the contact state, so a multi-contact hand or tactile fixture can be necessary. The right hand is the one that explains defects with the fewest blind spots.
Demonstration Interfaces Can Mismatch the Hand
Human demonstration interfaces make hand choice more complicated. UMI can collect teaching data without deploying the robot in the wild [12], but if the target hand is suction, many human contact strategies are not executable actions. DexUMI and DEXOP reduce the gap between human hands and robot hands [8] [28], but manufacturing transfer still has to be validated task by task. A human's slight twist during insertion may become torque-aware rotation for a two-finger gripper or in-hand reorientation for a dexterous hand.
Paired demonstrations are therefore important. Run the same part, fixture, and defect taxonomy with a human hand, a teleoperation device, and the target robot hand. Do not compare only success. Compare the contact cue before failure, the recovery path, and the operator note that disappears under each interface. This test turns hand choice into an embodiment-alignment problem that precedes policy architecture.
Force-aware demonstration interfaces preserve contact information that pose-only collection can discard, while adding hardware and calibration burden. UMI-FT addresses force-aware in-the-wild demonstration within three single-arm tasks [11], CoinFT offers one compact six-axis sensing implementation [6], and DexForce studies force-informed action targets on a particular hand and set of contact-rich tasks [33]. These results do not establish that teleoperation has become unnecessary or that total data cost always falls. Tethers, sensor delamination risk, task-specific compliance tuning, and the cost of translating force intent to a target hand without equivalent sensors remain.
Success rate alone is insufficient for comparing demonstration interfaces. Move the same part through the same fixture with a human hand, the collection device, and the target robot, then identify where pose, force, or contact events disappear at each transformation. DexUMI proposes the human hand as a manipulation interface [8] ([#8]); DEXOP explores a direct-contact passive exoskeleton [28] ([#10]); and a separate Terry note explains the force-aware UMI-FT direction [11] ([#36]). Those links are reading guides, not substitutes for the primary papers.
Maintenance Logs Are Learning Data
The missing variable in many hand strategies is maintenance. Suction-cup replacement, jaw-pad wear, tactile-skin replacement, finger calibration, and cable-tension adjustment can all change policy performance. If those events are detached from episodes, a hardware drift problem can look like model regression. Replaceable tactile skins such as ReSkin and AnySkin help [4] [5], but replaceability also means replacement events must be recorded.
Manufacturers should not leave hand maintenance only in a separate maintenance system. It should connect to the learning log. Before and after a model release, the team should know whether a sensor was replaced, a calibration fixture changed, or a cleaning procedure was modified. Tactile sensing increases data richness, but it also exposes learning to drift and contamination. Without operating logs, a richer hand becomes richer noise.
Changing the Hand Changes the Data Contract
Changing the end-effector is not just a bill-of-materials change. It changes the episode schema, replay set, simulation asset, maintenance log, and operator training. Even when interfaces such as UMI and DROID make in-the-wild collection easier [12] [22], some demonstration signal disappears when the target robot hand contacts the world differently. Hand migration therefore needs paired data: the same task performed with a human hand, a teleoperation device, and the target gripper, with the missing signals measured explicitly.
A practical hand experiment should proceed in sequence. First, build a failure taxonomy with the simplest viable end-effector. Second, determine whether the expensive failures are vision blind spots or force blind spots. Third, decide whether to add sensing or change the hand. Fourth, compare replay results before and after the hand change under the same defect codes. This turns a dexterity debate into a process-loss calculation.
Manufacturing Cell Checkpoint
A hand-selection review needs at least four tables: expected contact by task family, sensor coverage by failure code, model impact by maintenance event, and replay results before and after hand changes. Ask vendors for episode export fields before asking for URDFs or CAD. Grip force, suction pressure, tactile image, calibration version, tool wear, and operator override must share a key, or root-cause analysis will be rebuilt later inside a data lake.
Evidence tiers should remain separate. Tactile-sensor papers provide method and benchmark evidence. Company pages from Figure AI or Physical Intelligence reveal generalist policy and humanoid product direction [16] [17]. For a manufacturing cell purchase, durability, cleaning, spare parts, calibration drift, and defect reduction are more direct evidence.
A hand vendor or humanoid vendor should be evaluated by the same rule. A product page may show dexterous manipulation, but a manufacturing buyer needs spare-part cycles, sensor replacement, cleaning validation, calibration time, and failure export. Even with strong policy evidence from pi0 or Diffusion Policy-style work [19] [17], the policy cannot explain process causes if the hand does not leave stable signals in the field.
Questions for a Hand-Change Approval Meeting
A hand change should be approved in an operating review, not in an engineering demo alone. The first question is which defect the new hand reduces. The second question is whether that defect exists because of the current sensor blind spot. The third question is how maintenance events from the new hand connect to model release. If those questions cannot be answered, the hand change is more likely to add complexity than value.
Production, quality, maintenance, data, and the robot vendor should all be present. Production evaluates cycle time and cleaning burden. Quality owns the defect taxonomy. Maintenance evaluates spare parts and calibration time. The data team evaluates the episode schema and replay-set change. The vendor must explain not only hand capability, but which tactile, force, and raw state fields can be exported. This structure separates "the five-finger hand is impressive" from "the five-finger hand reduces manufacturing loss."
| Approval dimension | Required comparative evidence | Hold condition | Decision record |
|---|---|---|---|
| Task and failure coverage | Defect codes and contact states distinguished or missed by the current and candidate hands | Neither hand makes the expensive failure observable | State the included and excluded tasks |
| Sensor observability and calibration | Failure resolution of touch, force, and vacuum signals; calibration residual and clock offset | Output meaning cannot be compared before and after a sensor change | Freeze the selected channels and calibration criterion |
| Human and robot collection cost | Operator demonstration time, robot rollout time, intervention and reset burden | The new hand only increases volume without covering the required failures better | Set a budget and stop rule for each collection source |
| Cycle and quality impact | Shadow comparison of cycle change, defects, rework, and operator takeover | Throughput loss outweighs quality value or the denominator is unclear | Record the approved process metric and observation period |
| Maintenance and replacement | Cleaning, wear, spare parts, calibration time, and equivalence after replacement | Replacement events are not connected to episodes | Specify the maintenance-history key and recalibration procedure |
| Replay and rollback evidence | Shared replay cases, comparable historical-data scope, and a tested return to the old hand | The new schema silently invalidates the historical failure suite | Freeze the replay-suite version and rollback condition |
| Final decision | Joint review of quality value, operating burden, and data value | Evidence has no owner or an unresolved hold condition remains | Record approve, restricted approve, hold, or reject with rationale |
Finally, run a shadow period with the old and new hands. On the same SKU and fixture, run limited cycles with the current hand and the candidate hand, then compare defects and operator takeovers. Even if the new hand leaves richer signals, its total value can be lower if cleaning time grows or calibration drift is frequent. Hand choice is a data decision that includes cell economics, not just model accuracy.
The cost of a hand change has two sides. An under-instrumented hand fails to explain defects and raises relearning cost. An over-instrumented hand can reduce uptime through maintenance and cleaning. The former is a data blind spot; the latter is operations drag. A manufacturer should compare both in the same unit. If a tactile hand reduces scratch defects but needs calibration every shift, defect savings and downtime cost have to be read together.
That comparison leads directly into the flywheel design in the next chapter. Once the hand is chosen, the dominant data sources also change. A suction cell depends heavily on robot-native pressure logs and QA rejects. A dexterous-hand cell depends more on human demonstrations and tactile replay. The conclusion of the hand chapter is therefore a collection plan, not a hardware list.
A final approval artifact is a hand-change scorecard. It should show the old hand, the candidate hand, the expected defect reduction, the added sensor fields, new maintenance events, and the rollback plan. If the new hand changes the episode schema, the team should mark which historical replay cases remain comparable and which must be rebuilt. That prevents a hand upgrade from silently invalidating the data history that made the upgrade look justified.
Open Questions and Failure Modes
Tactile representations are improving faster than long-horizon evidence on drift and cleaning. Human-hand demonstrations lose contact semantics when they move to robot hands. Dexterous hands increase both data richness and maintenance burden. Putting a five-finger hand into a task that a simple gripper can solve can create a larger operating problem than the learning problem it was meant to solve.
Good hand choice is not the most impressive hardware. It is the lowest-risk hardware that observes the failures the process actually pays for. If that criterion is lost, large-data manipulation can accumulate more sensors and more complex policies while still failing to explain defects.
Figure 2.4. End-effector mechanism, sensing, calibration, and acceptance loop. This diagram is not a performance claim; it summarizes the operational chain that carries capability evidence through replay and quality decisions into release authority. Source: author-created SVG.
What to Learn Next
Chapter 3, Data Flywheels explains how human demonstrations, robot-native rollouts, simulation, and failure logs become one flywheel. If the hand defines the observation coordinates, the flywheel turns those coordinates into a continuously improving operating asset.
References
- Shaw, Kenneth (2023). LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. arXiv.
- Lambeta, Mike (2020). DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation. arXiv.
- Lambeta, Mike (2024). Digitizing Touch with an Artificial Multimodal Fingertip. arXiv.
- Bhirangi, Raunaq (2021). ReSkin: versatile, replaceable, lasting tactile skins. arXiv.
- Bhirangi, Raunaq (2024). AnySkin: Plug-and-play Skin Sensing for Robotic Touch. arXiv.
- Choi, Hojung (2025). CoinFT: A Coin-Sized, Capacitive 6-Axis Force Torque Sensor for Robotic Applications. arXiv.
- Feng, Ruoxuan (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors. arXiv.
- Xu, Mengda (2025). DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation. arXiv. #8
- Fang, Hao-Shu (2025). DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation. arXiv. #10
- Si, Zilin (2025). ExoStart: Efficient learning for dexterous manipulation with sensorized exoskeleton demonstrations. arXiv.
- Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv. #36
- Ha, Huy (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv.
- Yu, Wenhao (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
- Huang, Yuhang (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv.
- Hao, Yaru (2025). TLA: Tactile-Language-Action Model for Contact-Rich Manipulation. arXiv.
- Figure AI (2026). Figure 03 + Helix 02: General-Purpose Humanoid System. Company product page.
- Physical Intelligence (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. Company research post.
- Zhao, Tony Z. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
- Chi, Cheng (2023). Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv.
- Mandlekar, Ajay (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
- O'Neill, Abby (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
- Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
- Toyota Research Institute (2024). Large Behavior Models for Robot Manipulation. Company technical post.
- AgiBot-World Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
- Bicchi, Antonio (2000). Hands for Dexterous Manipulation and Robust Grasping: A Difficult Road Toward Simplicity. IEEE Transactions on Robotics and Automation.
- Capsi-Morales, Patricia et al. (2020). Exploring the Role of Palm Concavity and Adaptability in Soft Synergistic Robotic Hands. IEEE Robotics and Automation Letters.
- Agarwal, Arpit et al. (2025). A Modularized Design Approach for GelSight Family of Vision-based Tactile Sensors. The International Journal of Robotics Research.
- Fang, Bin et al. (2025). Force Measurement Technology of Vision-Based Tactile Sensor. Advanced Intelligent Systems.
- Gupta, Harsh et al. (2025). Sensor-Invariant Tactile Representation. arXiv.
- Zhao, Zihang et al. (2025). Embedding High-Resolution Touch across Robotic Hands Enables Adaptive Human-Like Grasping. Nature Machine Intelligence.
- Almeida et al. (2025). The Role of Touch: Towards Optimal Tactile Sensing Distribution in Anthropomorphic Hands for Dexterous In-Hand Manipulation. arXiv.
- Zhang, Chi et al. (2025). UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer. arXiv.
- Chen, Claire et al. (2025). DexForce: Extracting Force-informed Actions from Kinesthetic Demonstrations for Dexterous Manipulation. arXiv.
- Felix Berkenkamp et al. (2017). Safe Model-based Reinforcement Learning with Stability Guarantees. NeurIPS.
- Tuomas Haarnoja et al. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. ICML 2018.
- Kurtland Chua et al. (2018). Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models. NeurIPS.
- Alex Ray et al. (2019). Benchmarking Safe Exploration in Deep Reinforcement Learning. arXiv preprint.
- Tianhe Yu et al. (2020). Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning. CoRL.
- Shaoxiong Wang et al. (2022). Tacto: A Fast, Flexible, and Open-Source Simulator for High-Resolution Vision-Based Tactile Sensors. IEEE Robotics and Automation Letters.
- Zhenyu Wei et al. (2024). D(R,O) Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping. IEEE International Conference on Robotics and Automation (ICRA).
- Qian Mao et al. (2024). Multimodal Tactile Sensing Fused with Vision for Dexterous Robotic Housekeeping. Nature Communications.