Part III: Deployment Strategy

Chapter 11: Hardware Co-Design — Building for Learnability

Written: 2026-07-22 Last updated: 2026-07-22

Overview

A robot hand is not only an actuator. It is a data interface. Two-finger grippers and suction are fast, robust, and simple, but they can hide contact state. Five-finger hands, gloves, and tactile skins can expose more interaction while adding calibration, durability, cleaning, and retargeting costs. LEAP Hand [1] and multimodal tactile fingertips [2] show that dexterous research hardware is becoming cheaper and more instrumented, yet laboratory build cost and factory total cost of ownership are not the same metric.

Hardware co-design should not maximize anthropomorphic appearance or degrees of freedom. It should expose enough state to diagnose task failure while keeping control, calibration, cleaning, replacement, and requalification affordable. The tactile and force-rich direction in [4], [5], [14], and [13] matters because it records slip, squeeze, insertion, and compliance signals that vision-only manipulation often misses. As sensors multiply, the design must also establish which signals change a quality decision and whether the same data distribution can be restored after service.

After reading this chapter... - Compare hand choices by cost, cycle time, failure observability, and maintenance burden. - Explain which signals are lost when gloves, exoskeletons, or egocentric video become robot execution data. - Identify manufacturing tasks that need tactile or force sensing and tasks that do not. - Co-design sensing, actuation, end effectors, and controllers around failure observability and safety boundaries. - Evaluate hardware claims through calibration, serviceability, physical replay, and total-cost evidence rather than policy demos alone.

Task-Hand-Data Matrix

Task class Favorable hand / end-effector Signals that must be logged Main risk
Aligned pick/place, simple bin picking Two-finger gripper, suction, custom jaw Grasp pose, vacuum/force threshold, miss reason, downstream reject High success rate can hide slip and surface-damage causes.
Tool use, regrasp, human-tool compatibility Five-finger hand, anthropomorphic hand, glove retargeting [1], [6] Fingertip contact, joint limits, retargeting error, tool pose More degrees of freedom raise calibration and maintenance burden.
Deformables, food, cable, cloth Compliant gripper, tactile skin, force-torque sensing [4], [3] Squeeze, slip, local deformation, material state Vision-only data misses contact changes just before failure.
Insertion, fastening, wiping, contact-rich assembly Force-aware hand, tactile fingertip, guarded motion [14], [13] Contact onset, normal/shear force, recovery path, fixture tolerance Sensor contamination and recalibration cycles can dominate policy performance.
LEAP Hand component topology, human-hand scale comparison, and object and tool grasps

Figure 11.1. LEAP Hand Figure 1 shows the direct-driven hand's fingertip, motor, spacer, bracket, and palm topology, a human-hand scale comparison, and several object and tool grasps. This is hardware-layout and demonstration evidence, not a measured result for cycle life, calibration drift, factory yield, or total cost of ownership. Source: Shaw, Agarwal, and Pathak (2023), arXiv:2309.06440, Figure 1, PDF p. 1; figure-only crop from the primary paper.

Mechanical Observability

Degrees of freedom and payload capture only part of hand choice. In manufacturing manipulation, the hand is the device that observes physical state. A suction cup's vacuum trace reveals seal failure and porous material. A two-finger gripper's motor current reveals closure and hard stops. A five-finger hand's fingertip sensors can reveal slip, contact order, and local deformation. The same successful action has different learning value depending on what the hand recorded.

Call this mechanical observability: how directly the end-effector exposes the cause of success or failure. Suction may look low-observability, but vacuum pressure and cup wear can be sufficient for some cells. A complex five-finger hand can be low-observability if tactile baseline, joint calibration, and pad wear are absent from the episode. Hardware sophistication and data usefulness are not the same thing.

Hardware sophistication creates production value only when it adds diagnosable task state without unacceptable maintenance and calibration burden [9], [26], [27]. High-resolution touch, slip detection, and palm–finger coordination studies show ways to expose contact state on the hands, objects, sensors, controllers, and denominators they tested. Those results do not by themselves establish lower total data cost or independent reproducibility under factory contamination, cleaning solvents, cable fatigue, and multi-shift maintenance.

LEAP Hand [1] shows how lower-cost dexterous platforms can reduce the cost of hand experimentation. A manufacturer still has to separate research-platform cost from production total cost. Spare parts, cleaning, calibration time, operator training, and downtime can dominate the purchase price. A cheap hand that needs constant baseline reset may create expensive data.

Choosing the Hand Is Choosing the Data

Two fingers versus five fingers is not a hierarchy. A good manufacturing hand records success and failure cheaply, repeatedly, and explainably. Suction is fast and strong but fragile on porous or contaminated surfaces. Two-finger grippers are robust but give up most in-hand reorientation. Five-finger hands can reuse human tools and workspaces, but if per-finger contact and calibration state are absent from the episode schema, the extra complexity becomes noise.

LEAP Hand [1] shows how a low-cost dexterous platform can accelerate policy learning and hardware iteration. In manufacturing, though, the key question is not hand price alone. It is whether the hand produces the same data after shift changes, pad wear, backlash, cable-tension drift, and fingertip contamination. If those changes are not logged, model degradation becomes a mystery rather than a diagnosable maintenance event.

Task partitioning also changes. Traditional automation often chooses end-effectors by cycle time and grip success. Large-data manipulation adds a third question: through which sensor channel does the failure become visible? If scratches dominate cost, force limits and tactile patches matter. If missed picks dominate, camera coverage and suction trace may matter more. If insertion jams dominate, wrist force-torque and fixture state are critical.

A hand-selection review should not begin with CAD files and gripper catalogs. It should begin with prior defects and operator interventions. Convert those into failure codes, then mark whether each code is visible in vision, force, tactile, motor current, vacuum, or downstream QA. That table often shows that some cells need only suction or two fingers while others cannot close the data loop without a dexterous hand or custom compliant tool.

When the task and failure signals are structured around a small action space and passive compliance, a simple gripper or underactuated hand can provide better repetition and diagnosability than a high-degree-of-freedom hand [28], [29], [30]. The Pisa/IIT SoftHand lineage demonstrates the design logic of coupling many mechanical degrees of freedom to a small number of actuators and synergies so that mechanics absorb misalignment. Its grasp and in-hand results remain bounded to the reported hands and object suites; they do not rank all independently actuated hands under modern learned control.

Co-Designing Actuation and Control

Two hands with the same shape can produce different data when their actuators and transmissions differ. High-ratio electric drives can be position-repeatable but less backdrivable. Tendons and series elasticity absorb impacts while introducing tension, friction, and wear as latent states. Pneumatic or soft actuators can distribute contact but require pressure-to-shape calibration and may depend on temperature. Episodes should therefore include motor current, tendon tension, pressure, thermal state, saturation, and controller mode as well as commanded position.

The low-level controller is not plumbing outside the policy. Changing impedance, force limits, guarded motion, or slip response changes the contact outcome of the same high-level action. If a dataset retains only the action and omits controller gains and safety clamps, failure cannot be attributed between policy and control after a hardware revision. A co-design review should put policy update rate, servo rate, latency budget, sensor bandwidth, saturation, and emergency-stop authority in the same table.

More degrees of freedom do not imply more safety authority. A dexterous policy may choose contact order, but an independent safety layer should still constrain collision energy and pinch points. Conversely, a force limit that is too conservative can turn normal insertion and regrasp into failure. Hardware and controller must be qualified together to find the boundary between safety and learnability.

Co-design layer State that must remain in data Representative failure Safety/learning boundary
Actuator/transmission Current, pressure, tension, temperature, saturation Stall, backlash, tendon slip, degradation Torque/pressure limits and sustainable duty cycle
End effector Jaw/finger pose, pad or cup lot, tool ID Slip, crush, seal leak, tool loss Contact energy and product-damage limits
Sensing Raw/filtered signal, baseline, calibration epoch Drift, saturation, occlusion, contamination Reduce policy authority when sensor health is low
Controller Gain, mode, update rate, clamp, latency Oscillation, jam, delayed recovery Separate learned-action authority from certified safety authority
ORCA hand topology with tactile sensing, compliant joints, tendon tensioning, cooling, and author-reported deployment panels

Figure 11.2. ORCA Figure 1 places tactile sensors and silicone skin alongside a 17-DoF tendon-driven hand, 1-DoF wrist, cooling fans, overload-release joints, and ratchet-spool retensioning. Panels B and C are labeled by the authors as a 7+ hour imitation-learning deployment and 2,000+ cycles; these single-system demonstrations are not independent evidence of maintenance life, calibration stability, or production performance. Source: Christoph et al. (2025), arXiv:2504.04259v2, Figure 1, PDF p. 1, CC BY 4.0; figure-only crop from the primary paper.

Gloves and Exoskeletons Are Translators

Turning human hand data into robot execution data is attractive for large-data manipulation. UMI [15] collects human demonstrations outside the target robot setting; DexUMI [6] and DEXOP [7] make the human hand a more direct interface for dexterous manipulation. ExoStart [8] reinforces the idea that a small amount of robot data can be amplified by richer human data.

But gloves and exoskeletons are translators, not ground truth. Human skin sensation, force modulation, finger compliance, and arm-hand coordination do not map cleanly to robot joints and sensor layouts. Retargeting error, contact mismatch, operator strategy, and robot joint limits must be stored with the episode. A demonstration that a human performed well becomes a deployable robot trace only when the translation loss is measured.

DexUMI and DEXOP are important because they make that translation loss more explicit [6], [7]. They preserve more of joint, contact, and interface structure than video alone. Manufacturing tasks still need task-specific checks. Tool insertion may depend more on wrist force and fixture tolerance; cable routing may depend more on fingertip contact order.

Shared or aligned sensing between human collection and robot execution can reduce part of the embodiment gap, but transfer remains conditional on hand geometry, sensor calibration, camera layout, and action mapping [31], [32], [22]. RealDexUMI's use of the same lightweight hand and vision/tactile modules during collection and deployment shows how to reduce end-effector mismatch. Its reported average across eight tasks and three embodiments is bounded to that setup; it does not remove whole-arm dynamics or different camera geometry from the gap.

Wearable collection hardware must be evaluated for operator comfort and fatigue, sensor drift, hand–object occlusion, haptic fidelity, and repeated calibration as well as demonstrations per hour [7], [6], [33]. DexUMI's reported hand-platform average belongs to its exoskeleton, inpainting method, and task suite. Constraining human motion to robot-feasible kinematics can simplify retargeting while changing natural strategy and felt force; ergonomic cost under repeated long sessions remains a separate qualification problem.

Egocentric data sits at another layer. Human video can be broad and cheap, but tactile and force are mostly absent. Treat egocentric video as task-context and visual-prior data, then build release evidence from robot-side force and tactile replay. The more human data grows, the more disciplined the robot execution data must become.

Custom Hands Versus General Hands

Manufacturers often face a real trade-off between a general hand and a dedicated end-effector. A general hand may transfer across product changes and preserve human tools or workspaces. A dedicated hand is faster and more stable in a narrow task, and its failure causes are easier to isolate. From a large-data perspective, the trade-off is data reuse versus failure clarity.

A dedicated hand reduces action space. The same number of episodes covers the task faster, and root causes often cluster around jaw wear, suction leak, or fixture offset. The cost is redesign when the task changes and weaker reuse of human demonstration data. For fixed product families and high throughput, this simplicity may be the right data strategy.

A five-finger hand is the opposite. It can reuse more task structure, but it expands the failure space. Fingertip contact, joint limits, tactile calibration, object pose, and bimanual coordination all become variables. Figure- or Sanctuary-style humanoid-hand claims should be read through this balance. A human-like hand opens workspaces, but the manufacturer must carry the data and maintenance cost of that generality.

What Tactile and Force Stacks Unlock

Tactile and force sensing are not necessary for every task. They become close to essential when success depends on the state after contact: insertion, fastening, wiping, cable routing, deformable handling, and food portioning. Multimodal tactile fingertips [2] and small force-torque sensing [3] make contact richer; AnySkin and AnyTouch [4], [5] show how tactile skins and representations can make different sensors more learnable.

Policy architectures are moving in the same direction. ForceVLA [14] and Tactile-VLA [13] try to add physical contact state to language-action policies that would otherwise see only RGB. For manufacturers, the value is not the sensor novelty itself. It is whether slip onset, excessive force, fixture collision, or part misalignment is observed before the failure becomes a scrap or safety event.

Sensor placement matters. Fingertip tactile sensing is strong for grasp and in-hand motion. Wrist force-torque is often better for insertion, pushing, and total contact load. Tactile skin can reveal broader pressure distributions, but cleaning and wear become serious. Small force-torque sensors are compact, but interpretation depends on installation location [3]. The question is not whether the hand has touch. It is which failure code each sensor sees.

AnyTouch-style representation work shows how multiple tactile sensors may be bound into a shared learning space [5]. That has maintenance implications. If a sensor module or fingertip vendor changes but representation remains stable, dataset continuity improves. The claim remains incomplete, however, until it is tested under factory dirt, wear, temperature, and cleaning solvents.

Design choice Observability gained Operating boundary that must accompany it
Simple hand or suction Narrow, repeatable grasp and vacuum-state records Holdout replay for product and fixture changes
General multi-finger hand Richer contact-order and in-hand manipulation records Versioned per-finger calibration, wear, and recovery state
Tactile and force stack Earlier detection of slip, excess force, and contact transitions Recalibration after cleaning or replacement, linked to QA labels

Reading Hardware-Company Claims

Claims from Sunday, Eka, Genesis, Sanctuary, Figure, or similar hand and humanoid vendors should be read in three passes. First, look past degrees of freedom and payload to task evidence: which products, fixtures, and failure modes have been repeatedly tested. Second, inspect the sensor suite: "tactile" can mean normal force only, or it can include shear, slip, and contact images. Third, make sure maintenance data is connected to policy data.

When company material is thin, public research becomes the evidence floor. LEAP, DIGIT-style tactile fingertips, AnySkin, AnyTouch, DexUMI, DEXOP, ForceVLA, and Tactile-VLA tell us what questions a hardware claim must answer. A factory PoC should put the vendor's hand demo into this checklist and fill missing fields through instrumentation or contract terms.

This does not mean company claims should be dismissed. Hardware companies can expose durability, manufacturability, service coverage, and warranty terms faster than papers can. Manufacturers should read papers and company claims as complementary evidence. Papers show what is physically and algorithmically possible; companies must show that the system survives production maintenance.

Field Maintenance and Calibration

The more complex the hand, the less separable learning and maintenance become. Fingertip skins wear, sensors get dirty, joint zeros drift, and cable tension changes. If those changes live only in maintenance logs and not in episodes, data scientists will see policy drift while field engineers see sensor problems. The two views have to be joined.

The practical output of hardware co-design is therefore a calibration protocol, not only a hand drawing. The manufacturer needs to know when tactile baselines are reset, which replay set is run after cleaning, whether data before and after sensor replacement can be mixed, and whether spare-part lots change performance. A five-finger hand that cannot answer these questions is not advanced hardware for manufacturing; it is advanced uncertainty.

Replaceable tactile skins and modular sensors improve serviceability only when replacement triggers a defined recalibration and replay process [34], [35], [36]. ReSkin and modular GelSight designs make sensing surfaces or internal modules replaceable, while sensor-invariant representation work tries to reduce module differences in the learned space. Old and new data should not be merged automatically until physical tests show that zero shift, optical or magnetic variation, adhesion, and friction remain within the accepted envelope.

A replacement event should become a dataset boundary rather than end as a maintenance ticket. Record module serial and lot, pad material, mounting torque, firmware, and calibration artifact in episode metadata, then rerun known contact, slip, and over-force cases. Passing replay opens a new calibration epoch. Failure should send the team to installation and sensor health before policy retraining. Otherwise serviceable hardware creates frequent but silent distribution shifts.

From Simulation to Physical Qualification

CAD and simulation can prune hardware candidates quickly. Jaw geometry, reachable grasps, collision, actuator thermal load, cable routing, and sensor placement can be compared before repeated fabrication. Yet a simulator without sufficiently identified contact parameters and tactile response can reverse the ranking of two hands. Simulation is appropriate for design exploration and rare-case generation; it is not final evidence for physical friction, compliance, and wear.

Qualification is safer in three layers. Bench testing measures actuator saturation, backlash, force limits, sensor noise, and calibration repeatability. Instrumented-cell testing varies product lot, fixture offset, speed, and contamination while checking whether failure codes remain observable. Limited production then adds shifts, cleaning, maintenance, and operator handoff, and reads throughput, defects, and stop tails together. A demo becomes release evidence only when every layer has a pass criterion and a stated sample denominator.

Economics should follow the same layers. A sophisticated hand may save fixtures and tool changes while its spare modules, calibration station, skilled maintenance, or longer cycle time erase the gain. A simple gripper is cheap and fast but may increase SKU-specific tooling and changeover. Compare accepted units per available production hour, quality escapes, manual assist, tooling changes, sensor replacement, and validation time—not purchase price or a single success rate.

Qualification layer Conditions varied Evidence to pass What the layer cannot establish
Bench Load, speed, temperature, repeated calibration Saturation, noise, backlash, and baseline within limits Real product and operator interaction
Instrumented cell Lot, fixture offset, contamination, controller Failure-class detection and safe recovery Multi-shift maintenance and long-term wear
Limited production Shift, cleaning, maintenance, product mix Accepted units/hour, defects, stop tails, assists Universal transfer to another plant or embodiment
Change qualification Module, pad, firmware, policy version Lineage-aware replay and rollback rehearsal Absence of unobserved rare failures
UniTacHand human-to-robot tactile alignment and unified latent-space training architecture

Figure 11.3. UniTacHand Figure 1 maps paired human-glove and DexHand gesture/tactile inputs through morphology alignment and domain encoders into a shared latent space, supervised by pressure-UV reconstruction and contrastive objectives. The diagram specifies an alignment architecture and training losses; it does not measure whether replacement sensors preserve calibration, whether contamination is tolerated, or whether factory transfer remains stable. Source: Zhang et al. (2025b), arXiv:2512.21233v3, Figure 1, PDF p. 2; complete figure artwork from the primary paper.

Cell-Level Hand Selection Walkthrough

The first cell picks metal parts from an aligned tray and places them into an inspection fixture. A two-finger gripper or custom jaw may be better than a five-finger hand. The key failures are missed pick, scratch, and fixture misalignment. Jaw surface wear, grasp force, fixture contact, and downstream reject may be enough; fingertip tactile sensing may be less economical than wrist force or gripper current.

The second cell routes a cable into clips. It may still look like pick-and-place, but contact order and compliance dominate. A two-finger gripper may miss cable twist and partial insertion. DexUMI or DEXOP-style interfaces can capture human routing strategies, while robot-side tactile or force replay measures transfer loss [6], [7].

The third cell portions sauce or another viscous material. Human hand shape matters less than material state and tool contact. Force-torque, utensil pose, container geometry, and weight QA need to attach to the episode. Even if a Chef Robotics-style platform is considered, material-state logging and operator-correction export should be examined before hand shape [25].

The fourth cell uses a human tool in an existing assembly workspace. A five-finger hand may preserve tools and layout. It also requires tool pose, fingertip contact, grip transition, and failure recovery data. Figure- or Sanctuary-style humanoid hand claims are attractive here, but without tactile use and maintenance evidence, long-term production risk remains high [10].

These cases show that hand choice is not a catalog question. It is a failure-observability question. The right strategy is not to standardize one hand everywhere. It is to identify the minimum observability each failure code requires and choose the hand around that.

Hand decisions may also change over time. A manufacturer might start with a two-finger gripper to collect data quickly, then move to tactile fingertips when failures prove contact-rich. Conversely, a five-finger hand may be simplified if most failures turn out to be perception or fixture problems. A large-data strategy should allow hardware revision based on evidence.

Hardware revision must become dataset versioning. A new jaw, fingertip material, tactile baseline, or sensor module changes the data distribution. If that change is not recorded, old and new episodes become difficult to mix. The policy may appear to improve because the new hand hides a failure, or appear to degrade because a sensor replacement shifted the baseline.

Manufacturers should ask vendors for lifecycle data, not only task performance. Replacement cycles, calibration drift, cleaning procedures, spare-part lot variation, and sensor failure modes all matter. Papers show hand possibility; production requires lifecycle continuity.

Data collection changes with hand choice. A two-finger cell may benefit from many fast robot rollouts. A dexterous-hand cell may benefit from fewer high-quality demonstrations and sensor-rich replay. A glove cell depends on robot retargeting validation. A tactile-rich cell needs calibration and QA linkage before scale.

Ignoring those differences reverses the data strategy. Excessive glove capture on simple pick-place raises cost. Camera-only fleet rollout on dexterous assembly hides failures. The safer order is to define the data needed first, then choose the hand.

The last test is whether operators and maintenance teams can understand the hand. A complex hand that enters the line without field language turns every small abnormality into an "AI problem." If sensor meaning, wear modes, and calibration state are translated into operator and maintenance workflows, people leave better corrections and better failure labels.

This is especially important for tactile systems. A fingertip image, a shear estimate, or a force spike is only useful if the organization knows what action follows. Does the operator clean the fingertip, rerun calibration, reject the part, or tag a replay episode? Without this operational meaning, tactile data remains an engineering artifact rather than a production signal.

Manufacturers should therefore review hand vendors with both robotics and maintenance staff in the room. Robotics teams can judge policy compatibility; maintenance can judge serviceability; quality can judge label value; production can judge cycle-time disruption. The chosen hand is the one that survives all four views, not the one with the most impressive dexterity video.

The same review should include spare parts and consumables. Pads, skins, cables, fingertips, gloves, and sensors are not neutral replacements. A different material lot can change friction, compliance, tactile baseline, or cleaning behavior. The data schema should record those changes so the model team can separate policy error from hardware drift.

Lifecycle logs also decide whether imitation data remains usable. A glove or exoskeleton session may capture excellent human strategy, but it can become misleading if the robot hand cannot reproduce the same contact order after wear or recalibration. DexUMI, DEXOP, and ExoStart-style work make demonstration capture more practical; factory deployment still needs a robot-side replay gate after hardware changes.

For tactile systems, representation continuity is the hard issue. AnySkin, AnyTouch, DIGIT-style fingertips, and Tactile-VLA point toward transferable contact representations, but a plant has to ask how transfer is validated after cleaning, replacement, contamination, or temperature change. Without that validation, tactile data can make the dataset look richer while making release decisions less stable.

The business case should also be hand-specific. A five-finger hand can be justified when it preserves tools, fixtures, and human workspaces, but its maintenance burden must be priced into the cell. A simple gripper can be justified when it yields faster learning and lower downtime, but it may push complexity into fixtures or SKU-specific tooling. The data strategy should expose that tradeoff instead of hiding it inside a robot quote.

For that reason, the final hand review should include a negative recommendation. The team should name the tasks where the proposed hand should not be used, the failure modes it cannot observe, and the data that would be needed to change that decision. A bounded no is often more useful than a vague yes because it keeps hardware scale-up tied to evidence.

Manufacturing Cell Checkpoint

Checkpoint Two-finger / suction / dedicated hand Five-finger / general hand Glove or exoskeleton collection Tactile / force stack
Data unit Grasp pose, suction trace, miss reason Finger contact, joint state, calibration version Retargeting error, operator strategy, robot limit Contact onset, force/shear, slip, baseline
Best fit High repetition, simple pick/place, high throughput Tool use, regrasp, human-workspace compatibility Early process discovery, dexterous demonstrations Insertion, deformables, food, cable, wiping
Main cost Tooling changes and SKU drift Maintenance, cleaning, calibration Translation loss and consent Sensor wear, representation continuity
Release gate Fixture/SKU holdout replay Replay before and after hand calibration Robot-side validation replay Tactile/force failure replay

This checkpoint does not choose the hand by itself. It identifies missing information. A simple hand is better when it preserves the right failure signals. A dexterous hand is justified when the task truly needs contact order, tool compatibility, or regrasp. A sensor without a release gate adds storage cost, not production learning.

Open Questions and Failure Modes

First, anthropomorphic hardware can be mistaken for human-equivalent manipulation. Second, glove data can be treated as robot-ready data without measuring translation loss. Third, tactile sensors can be installed without being connected to QA labels, leaving slip, crush, and misalignment outside the learning target. Fourth, calibration and cleaning can be separated from operating data, making post-deployment performance changes hard to explain.

The open questions are about standardization. If general dexterous hands become cheap and durable enough, how much advantage remains for dedicated end-effectors? If tactile representations transfer well across sensors, can maintenance burden fall? Or will contamination and cleaning keep breaking representation continuity? These questions cannot be answered by policy papers alone; they require field maintenance logs.

What to Learn Next

The final chapter ties these choices to ownership. Whether a manufacturer buys a model, a hand, a dataset, or a vendor-managed workcell, the factory flywheel is not an owned asset unless process truth and evaluation authority remain inside the manufacturing organization.

References

  1. Shaw, Kenneth (2023). LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. arXiv.
  2. Lambeta, Mike (2024). Digitizing Touch with an Artificial Multimodal Fingertip. arXiv.
  3. Choi, Hojung (2025). CoinFT: A Coin-Sized, Capacitive 6-Axis Force Torque Sensor for Robotic Applications. arXiv.
  4. Bhirangi, Raunaq (2024). AnySkin: Plug-and-play Skin Sensing for Robotic Touch. arXiv.
  5. Feng, Ruoxuan (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-Tactile Sensors. arXiv.
  6. Xu, Mengda (2025). DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation. arXiv. #8
  7. Fang, Hao-Shu (2025). DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation. arXiv. #10
  8. Si, Zilin (2025). ExoStart: From 10 Exoskeleton Demos to Dexterous Robot Manipulation. arXiv.
  9. Zhao, C. et al. (2025a). Universal Slip Detection of Robotic Hand with Tactile Sensing. Frontiers in Neurorobotics.
  10. Figure AI (2026). Figure 03 + Helix 02: General-Purpose Humanoid System. Company product page.
  11. Figure AI (2025). Helix: A Vision-Language-Action Model for Generalist Humanoid Control. Company announcement.
  12. Hao, Yaru (2025). Tactile-Language-Action Model for Contact-Rich Manipulation. arXiv.
  13. Huang, Yuhang (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv.
  14. Yu, Jiawen (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
  15. Ha, Huy (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv.
  16. Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv.
  17. Zhao, Tony Z. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
  18. Chi, Cheng (2023). Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv.
  19. Physical Intelligence (2025). OpenPI: Open Source Robot Policy Stack. GitHub.
  20. Black, Kevin (2024). pi0: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
  21. AgiBot-World-Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
  22. Xu, Chaoyi et al. (2026). RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning. arXiv.
  23. Generalist AI (2025). GEN-0 Robot Foundation Model. Company page.
  24. Skild AI (2024). General-Purpose Robot Brain. Company page.
  25. Chef Robotics (2025). ChefOS: AI Robotics Platform for Food Manufacturing. Company product page.
  26. Zhao, Zihang et al. (2025b). Embedding High-Resolution Touch across Robotic Hands Enables Adaptive Human-like Grasping. Nature Machine Intelligence.
  27. Zhang, Ningbin et al. (2025a). Soft Robotic Hand with Tactile Palm-Finger Coordination. Nature Communications.
  28. Bicchi, Antonio (2000). Hands for Dexterous Manipulation and Robust Grasping: A Difficult Road Toward Simplicity. IEEE Transactions on Robotics and Automation.
  29. Catalano, Manuel G. et al. (2014). Adaptive Synergies for the Design and Control of the Pisa/IIT SoftHand. The International Journal of Robotics Research.
  30. Della Santina, Cosimo et al. (2018). Toward Dexterous Manipulation with Augmented Adaptive Synergies: The Pisa/IIT SoftHand 2. IEEE Transactions on Robotics.
  31. Zhang, Chi et al. (2025b). UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer. arXiv preprint. #16
  32. Yin, Jessica et al. (2025). OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer. arXiv preprint. #18
  33. Fang, Hongjie et al. (2024). AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild. arXiv.
  34. Agarwal, Arpit et al. (2025). A Modularized Design Approach for GelSight Family of Vision-Based Tactile Sensors. The International Journal of Robotics Research.
  35. Gupta, Harsh et al. (2025). Sensor-Invariant Tactile Representation. arXiv.
  36. Bhirangi, Raunaq et al. (2021). ReSkin: Versatile, Replaceable, Lasting Tactile Skins. arXiv.
  37. Christoph, Clemens C. et al. (2025). ORCA: An Open-Source, Reliable, Cost-Effective, Anthropomorphic Robotic Hand for Uninterrupted Dexterous Task Learning. arXiv.