Chapter 12: Manufacturing Data Sovereignty — What Operators Must Own
Overview
The conclusion of large-data manipulation is not that every manufacturer should build its own foundation model. Vendors can provide foundation models, humanoid stacks, and workcell software. Manufacturers still have to retain process truth, evaluation harnesses, failure taxonomies, calibration lineage, data rights, and update authority. Open X-Embodiment, DROID, and AgiBot World show how quickly robot data scale is growing [1], [2], [3]. Factory performance improves only when that data is tied to quality decisions and operating governance.
Buy versus build is therefore not decided by demo performance. The decision turns on who defines the episode schema, freezes the replay set, manages QA labels and worker consent, approves updates, and can roll back a release. Data sovereignty here does not mean copying every file onto an internal server. It means durable control sufficient to explain the process, audit changes, stop deployment, and recover operations when suppliers change.
After reading this chapter... - Separate vendor-owned models from manufacturer-owned process truth. - Define export, replay, and rollback rights that belong in robot-data contracts. - Distinguish ownership, access, and use rights for raw telemetry, derived labels, schemas, calibration, and model artifacts. - Put provenance, retention, export, and portability requirements into PoC and procurement gates. - Place public datasets, company stacks, and factory telemetry in one decision system while separating simulation from real-cell release authority.
Six Assets Manufacturers Must Own
| Asset | Why the manufacturer must own it | Minimum artifact |
|---|---|---|
| Process truth | Vendors see tasks; manufacturers see defects, rework, scrap, and downtime. | Task taxonomy, defect codebook, cell acceptance criteria |
| Episode schema | Learning fails not only from missing data but from disconnected data. | Shared keys for observation, action, contact, controller mode, QA result, operator note |
| Replay / evaluation harness | Updates must be tested against prior failures before deployment. | Frozen replay set, holdout lots, safety-critical scenarios |
| Safety envelope / rollback | Model improvement is deployable only inside a safety case. | Force limits, guarded motion, stop rules, rollback playbook |
| Data rights and consent | Worker video and process IP are learning assets and governance risks. | Data-use policy, consent flow, retention/export rules |
| Vendor exit plan | Process knowledge has to survive vendor replacement. | Export format, model-independent logs, migration test |
Three Layers of Ownership
The ownership question is broader than whether the manufacturer trains the model. The first layer is process meaning. Which attempts count as the same task? Which defects belong to the same family? Which recoveries are allowed? This authority should not move to the vendor. Vendors can optimize robot behavior, but product quality and customer responsibility stay with the manufacturer.
The second layer is data structure. Episode schema, failure taxonomy, replay sets, and QA traces live here. Even when a vendor model is used, the manufacturer should know how its attempts are stored, which labels attach to them, and which update gates they feed. Large datasets such as Open X-Embodiment and DROID provide broad priors, but they do not replace plant-specific QA-linked schemas [1], [2].
The third layer is operating authority. Safety envelopes, release approval, rollback, worker consent, retention, and vendor exit live here. This layer is as much organization design as model evaluation. A strong policy that updates outside the safety case is not an asset; it is an operating risk.
Manufacturers need durable control of task definitions, episode schemas, frozen replay sets, quality traces, update approval, rollback, and export rights across vendor and model changes [25], [26], [27]. Runtime monitoring, real-world RL challenge analysis, and runtime assurance show—on their own robots and environments—why partial observability, nonstationarity, safety controllers, and human intervention belong in the operating structure. They do not legally guarantee a governance design or establish production readiness across manipulation cells. Durable control must be implemented against the plant's task, sensing, control, and quality denominator.
Control Rights and Provenance by Asset
"Ownership" is ambiguous if it refers only to legal title over files. For raw telemetry, specify who may collect, inspect, retain, and delete. For derived labels, specify who defines, changes, and approves them. For model artifacts, specify who can deploy, pause, export, and roll back. A cloud-managed system can still give a customer control of schemas and replay results; an on-premises file can still be unusable if only the vendor has its decoder and calibration.
| Asset layer | Required provenance | Minimum control right | Portability test |
|---|---|---|---|
| Raw telemetry | Cell/robot/sensor ID, clock, firmware, calibration epoch | Purpose-scoped access, retention, deletion or preservation hold | Reconstruct a time-aligned episode with an independent tool. |
| Derived labels | Label schema/version, annotator or inspector, confidence | Defect mapping and relabel approval | Rebuild the quality trace without the vendor dashboard. |
| Schema and calibration | Field dictionary, units, transform, baseline artifact | Change approval and backward mapping | Verify that a new collector remains compatible with old episodes. |
| Model and controller | Weights or API version, preprocessing, gains, safety clamp | Release, pause, rollback, and result audit | Compare commands and constraints on frozen inputs. |
| Replay and report | Inclusion reason, expected outcome, denominator, sign-off | Test execution and pass/fail criteria | Execute the same failure packet on another stack. |
The provenance chain should run raw → derived → training set → model/controller → deployment → quality outcome. If a link is missing, improvement becomes hard to explain. A changed reject code mixed with old labels, or a sensor replacement without a calibration epoch, prevents the team from separating model change from process change. Governance is not indefinite storage of everything; it is preservation of the lineage used for decisions for a reproducible period.
Figure 12.1. W3C PROV-DM Figure 1 defines core provenance relationships among Entity, Activity, and Agent, including use, generation, derivation, attribution, and association. This informative standards diagram supplies an interchange vocabulary; it does not show that a plant captured complete lineage, applied an adequate retention schedule, or connected an artifact to the correct QA decision. Source: W3C (2013), PROV-DM §2.1, Figure 1; official figure asset, W3C Document License.
The Boundary Between Vendor Models and Process Truth
It is natural for vendors to have stronger model priors. Open X-Embodiment, DROID, and AgiBot World provide a scale that most manufacturers cannot reproduce alone [1], [2], [3]. PI, Covariant, Dexterity, Chef Robotics, and Figure also accumulate their own deployment and data advantages [7], [8], [9], [10], [11].
Process truth cannot be outsourced in the same way. Only the manufacturer can define whether a torque spike is an acceptable press-fit, whether a smear is a food QA reject, which operator correction is safety-critical, or which rework later becomes a customer defect. A vendor model can learn from that truth, but the authority to define, approve, and audit it must stay with the manufacturing organization.
Process truth is often weakly documented. Skilled workers may say a part "feels wrong." Quality teams may record only inspection codes. Process engineers may manage fixture tolerance in spreadsheets. Large-data manipulation requires those fragments to attach to robot episodes. If that work is skipped, a foundation model does not learn tacit knowledge; it amplifies gaps in tacit knowledge.
A healthy boundary looks like this: the vendor provides broad priors, model tooling, deployment software, and fleet learning; the manufacturer provides task definitions, defect semantics, acceptance criteria, and line-specific replay. Both sides may look at the same data, but their authority differs. The vendor asks whether behavior improved for the model. The manufacturer asks whether behavior is acceptable for the process.
Rewriting Buy Versus Build
Build does not have to mean training the foundation model from scratch. At minimum, the manufacturer should build the evaluation harness and data governance layer. An open stack or internal baseline is useful, but the more important capability is comparing vendor candidates on the same process trace. Without that capability, buying becomes a decision based on demos and trust.
Buy does not mean transferring responsibility. A vendor-managed stack still needs manufacturer-defined acceptance tests, update windows, rollback thresholds, data export, and worker consent. With production-flywheel vendors such as Covariant, Dexterity, and Chef Robotics, contracts should separate whose model gets better from whose process gets better [8], [9], [10].
The buy/build question becomes three questions. Which layer can be delegated without losing process meaning? Which data fields must remain internal for root-cause analysis and audit? Will replay sets and failure taxonomies survive vendor replacement? Answering these lets manufacturers use external models without giving up the core learning asset.
| Boundary | What a vendor may provide | Authority the manufacturer retains |
|---|---|---|
| Model and tools | General prior plus training and deployment software | Process-specific acceptance criteria and comparative evaluation |
| Data loop | Fleet learning and candidate updates | Failure taxonomy, replay set, and quality decisions |
| Operating change | Release bundle and support procedure | Approval, pause, rollback, and vendor-transition testing |
Evaluation Harnesses and Replay Sets
The risk in large manipulation models is that average performance improves while specific factory failures disappear under the aggregate. Simulation evaluation keeps improving [17], but it does not replace holdout lots, dirty sensors, worn grippers, operator handoffs, and downstream inspection. Manufacturers need factory replay sets that are separate from vendor benchmarks.
A good replay set is not a gallery of successes. It includes historical scrap causes, near misses, slow cycles, emergency stops, hard-to-see misalignments, and human recoveries. Before any update, the system should pass that replay set, then move through limited-cell shadow or canary deployment. Whether the stack is open like OpenPI or closed and vendor-managed, update approval has to connect to the manufacturer's quality system.
Replay sets are not static artifacts. Product changes, supplier changes, tool wear, sensor replacement, and recipe updates should create versions. Some failures may become less important; new defect codes may become critical. The key is that a release replay set cannot be changed casually for each update. Sets used for learning and sets used for release decisions must stay separate.
Simulation belongs inside this structure. X-Sim and the study that crosses the human–robot embodiment gap with one human demonstration and sim-to-real reinforcement learning examine different paths from human demonstrations through simulation to candidate robot policies, while SIMPLER evaluates simulation–real policy-behavior correlation in covered setups [28], [29], [17]. NVIDIA's announced Omniverse libraries and Blender workflow are first-party platform context for scene authoring and simulation-ready validation [14]. The announcement supplies no factory-reality-equivalence or policy-readiness metric, so capability evidence must not be promoted to assurance evidence.
Simulation, learned world models, and synthetic data can propose candidate policies and stress cases, but they cannot confer release authority without versioned real-cell evidence [28], [29], [17]. Each sim-to-real or correlation result is bounded to its task, embodiment, scene reconstruction, controller, and evaluation denominator. Simulation replay should complement, not replace, a final gate containing holdout lots, contamination, wear, operator handoff, and downstream inspection.
Human Data, Worker Consent, and Process IP
Large-data manipulation will keep using human data. UMI, UMI-FT, and offline demonstration research show how human demonstrations and human-in-the-loop learning can reduce the robot-data bottleneck [4], [5], [18]. Factory video and manipulation traces, however, contain worker identity, skilled practice, process know-how, and customer product information. They are learning assets and governance objects at the same time.
Manufacturers should separate three categories. First, raw data retained for internal process improvement. Second, anonymized or aggregated data that may be used for vendor model improvement. Third, data reserved for audit and never used for external training. Without these categories, a data flywheel can turn into labor conflict, customer-IP exposure, or vendor lock-in.
Worker consent is not only a legal document. A good consent flow explains what video is collected, which fields can leave the plant, and which data will not be used for individual performance management. If operators believe their corrections become productivity surveillance, human-in-the-loop learning will fail. Learning data and performance-management data should be separated.
Process IP needs the same care. Robot traces contain product geometry, assembly order, defect patterns, and skilled workarounds. Those are valuable to the vendor and strategically sensitive to the manufacturer. Contracts should separate raw video, masked video, trajectories, derived embeddings, and aggregate metrics. Blocking all sharing weakens vendor learning; sharing everything externalizes process knowledge.
Access and retention should differ by data type. Live raw-video access may open narrowly during incident investigation, while event codes and quality traces may remain longer for trend analysis. Training snapshots and release replay should remain long enough to reproduce why a model was approved. This chapter cannot prescribe a universal duration: the schedule must be approved against product, quality, labor, customer, and organization-specific obligations.
Export is not a one-time file dump. Schema, units, clocks, transforms, missing-data semantics, calibration artifacts, and label dictionaries have to travel with the records. If an export package is not read regularly in an independent environment, format errors will first appear during vendor exit. Governance overlaps with generic security and compliance but is not identical to either. Security protects access and systems; compliance addresses applicable obligations; this governance layer records which evidence made an operating decision and whether that decision can be reversed.
Figure 12.2. IDTA's Asset Administration Shell specification Figure 57 places AutomationML and OPC UA representations above XML, JSON, and RDF payloads, concept-description standards, communication, and type/instance lifecycle layers. This is a first-party standards-body mapping diagram, not proof that a vendor export is conformant, round-trippable, replayable, or complete for clocks, calibration, and missing-data semantics. Source: IDTA (2025), AAS Specification Part 1 v3.0.2, Figure 57, PDF p. 110; figure-only crop.
Safety Boundaries and the Operations Layer
The stronger the policy model, the more important the operations layer becomes. A stronger model proposes actions in more situations, while the factory still requires force limits, speed limits, safety zones, fixture interlocks, and line-stop rules. ForceVLA and Tactile-VLA point toward richer contact-aware policies [21], [22], but richer signals should still execute inside explicit safety boundaries.
The operations layer also governs updates. A performance change may come from model weights, prompt or task description, controller parameters, gripper-pad replacement, sensor cleaning, or fixture drift. Rollback is only possible if the system records what changed. Otherwise the organization is left with a vague diagnosis that "the AI got worse."
A release gate should compare real and simulated trajectories, contact events, failure types, timing, and quality outcomes under versioned calibration [15], [30], [17]. A simulator that matches visual action trends does not guarantee hidden contact force, product damage, or sensor drift. A physical mismatch should update the parameter and sensor-model discrepancy record rather than simply discard simulation. This comparison does not turn correlation in a source setup into general safety equivalence.
The operations layer should also treat human handoff as a normal state, not only a failure. In early production flywheels, operator correction may be the best label. If the system logs why a human took over, how the human recovered, and whether QA passed afterward, recovery policy can improve. If handoff is hidden, success rates can look good while true labor substitution remains low.
This layer also intersects with IT/OT security. A robot model update is more dangerous than an ordinary software update because physical behavior changes. The plant should record update artifact, release note, changed fields, rollback hash, and approver. The more cloud-managed the vendor system is, the more explicit the audit trail must be.
Safety and functional-assurance evidence must remain versioned and independent of a model vendor's performance narrative [27]. SOTER demonstrates a runtime-assurance design in which a declared safety specification and certified fallback controller constrain a high-performance controller under stated assumptions. Its case study uses drones and does not certify manufacturing manipulation. A plant must be able to audit safety limits, fallback, tests, approver, calibration, and rollback bundle outside the vendor dashboard.
| Release evidence | Comparison unit | Approval question | Artifact to retain |
|---|---|---|---|
| Simulation stress | Scenario, seed, physics/sensor version | Are expected failures and action trends reproduced? | Scene, parameters, trajectory, mismatch note |
| Real-cell replay | Lot, fixture, calibration, controller version | Are contact, timing, and quality outcomes within limits? | Raw/derived trace, QA result, denominator |
| Safety assurance | Limits, fallback, interlock, human handoff | Does an independent layer restrict authority under hazard? | Safety specification, test report, approver |
| Release/rollback | Model, preprocessing, gains, hardware bundle | What is promoted and what must roll back together? | Signed manifest, rollback rehearsal |
A 2026-2030 Execution Roadmap
The first step is not to pick the easiest task; it is to pick an instrumentable cell. A good first PoC has visible failure causes, fast quality feedback, logged operator intervention, and a vendor willing to agree on data export. The manufacturer should define the task taxonomy and episode schema before choosing the final model candidate.
Second, compare model options in parallel. Open stacks, production vendors, and humanoid vendors should run against the same replay set. Success rate should be evaluated alongside cycle-time tails, recovery success, operator assist, and downstream reject. Third, validate update governance during limited deployment: record which updates were approved and which replay failures blocked release. Fourth, scale by replicating failure taxonomy and data rights across cells, not merely by adding more robots.
The risky transition is pilot to scale-up. During a pilot, expert engineers stand nearby, vendor teams respond quickly, and operators understand the experiment. At scale, that protection disappears. A pilot should end only when operations are ready: the field team can update replay sets, quality understands rollback criteria, maintenance records sensor drift, and IT can audit data export.
The 2026-2030 roadmap should therefore include an ownership roadmap separate from the model roadmap. In 2026, build the instrumented cell and data contract. In 2027, compare vendors and internal baselines against the same replay. In 2028, replicate update governance across lines. By 2029 and beyond, keep a process data layer that survives vendor or model-family change.
Organization Operating Model
The assets manufacturers must own do not belong to one team. Process engineering owns task taxonomy and fixture variables. Quality owns defect codes and acceptance criteria. Production owns handoff patterns and cycle-time pressure. Safety owns force limits and line-stop rules. IT/security owns data movement and access control. A large-data manipulation project connects these into one episode schema.
A common early failure is letting the robot team choose the model and gripper alone. Quality codes then get attached late, maintenance logs stay outside the dataset, and worker consent appears near deployment. Bringing everyone into a large committee from day one is too slow, though. A small core team with named approvers for each ownership asset is more practical.
Operations meetings should also change. A meeting that reviews only weekly success rate is insufficient. Replay failures, new defect codes, operator-correction patterns, sensor drift, maintenance events, and vendor update requests should be reviewed together. Only then can the organization distinguish model change from process change.
Procurement matters as well. A robot vendor contract looks like equipment purchasing, but it is also a data and update-rights contract. If procurement evaluates only price and SLA, later engineering teams may struggle to recover process knowledge. Episode schema, replay rights, rollback, audit trail, and vendor exit format belong in the purchasing checklist.
A vendor exit plan requires model-independent logs, interpretable data export, frozen replay assets, and a tested rollback before operational dependency becomes irreversible [25], [31], [32]. Autonomous reset–rollout–verify and automated digital-twin workflows illustrate candidate update-loop components, but robot idle time, verification infrastructure, scene artifacts, and nondeterminism remain. No individual preprint or project claim guarantees manufacturing portability. The meaningful evidence is an exit drill that reconstructs episodes, reruns the failure pack, and starts a safe prior release without vendor staff.
Procurement should place gates at three points. An RFP requires export fields, API/format versioning, and notice of support termination. PoC exit reads a sample export with an independent parser and executes replay and rollback. Renewal reconstructs recent episodes, defect mappings, and release history without vendor staff. These are not declarations of distrust; they are operating tests that clarify responsibility for both parties.
| Procurement gate | Prior requirement | Test to execute directly | Decision on failure |
|---|---|---|---|
| RFP / design | Schema and API versioning, calibration export, end-of-support notice | Read a sample episode and field dictionary with an independent parser. | Hold the PoC until the interface is corrected. |
| PoC exit | Replay/rollback rights, access and retention matrix | Rerun the frozen failure pack and restore the prior bundle. | Do not promote to limited deployment. |
| Renewal / scale | Migration format, recent provenance, exit assistance | Reconstruct episodes, defects, and release history without vendor staff. | Treat dependency growth as contractual and technical risk. |
Figure 12.3. NIST AI RMF Figure 5 organizes Map, Measure, and Manage around the cross-cutting Govern function. It is a voluntary, cross-sector governance architecture; it is not a manufacturing audit result, replay or retention prescription, functional-safety certification, or evidence that a particular release is production-ready. Source: NIST (2023), AI RMF 1.0, Figure 5, PDF p. 25 (document p. 20); figure-only crop, republished courtesy of NIST.
The executive metric should be the number of learnable cells, not only the number of robots. A learnable cell has task definition, failure taxonomy, replay set, QA trace, safety case, and data governance connected. When that unit scales, the manufacturer turns vendor models into its own process asset.
The operating model also improves vendor relationships. A manufacturer with replay sets and defect taxonomies can ask for specific improvements: reduce over-force failures on fixture B that lead to downstream scratch rejects. That is more useful to vendors than vague dissatisfaction. Clear failure packets make joint improvement easier.
When ownership is weak, vendor management becomes a trust argument. The line says the robot got worse, the vendor says average KPIs improved, and quality sees downstream defects separately. Without a shared episode key and replay set, these claims cannot be reconciled. Ownership is not control for its own sake. It is the shared language for solving problems.
Training is part of ownership. Process engineers do not need to master VLA architecture, but they should read replay failures and update notes. Quality engineers do not need to write controllers, but they should know which episode fields connect to defect codes. Maintenance technicians should understand how sensor drift appears in policy metrics. Without this literacy, interpretation is outsourced.
Ownership becomes more important over time. In the first deployment, vendor experts may solve many problems. Two years later, product mix, suppliers, shifts, sensor lots, and safety rules will have changed. The durable asset is not the original model name. It is the manufacturer's process trace and replay discipline.
The organization should budget for that discipline explicitly. Replay maintenance, schema updates, consent audits, and vendor-export tests are recurring work, not launch tasks. If they are unfunded, they disappear after the pilot. The result is familiar: the first cell works under expert attention, then later cells accumulate silent exceptions.
Manufacturers should also measure ownership maturity. One simple metric is the percentage of robot failures that can be traced from episode to defect code to corrective action. Another is the percentage of model or hardware updates that pass through a documented replay gate. These metrics are less glamorous than task success, but they predict whether automation becomes a learning process.
The final institutional lesson is that ownership is not anti-vendor. Vendors improve faster when customers provide precise, auditable failure packets. Customers negotiate better when they know which process fields matter. The strongest deployments will likely be partnerships where vendors own broad model scale and manufacturers own process truth.
That partnership needs a governance rhythm. Before each model, controller, fixture, or hand update, the team should identify the changed artifact, choose the replay pack, name the approver, and record the rollback point. After the update, it should compare not only success rate but also cycle-time tail, defect-code shift, manual assist, safety stop, and maintenance note. This makes update approval a manufacturing process rather than a software habit.
Ownership should also be tested during procurement renewal. A vendor may be performing well, but the manufacturer should still run an export drill: can it reconstruct episodes, replay sets, defect mappings, operator-correction logs, and update history without vendor staff? If the answer is no, the renewal is also a dependency increase. If the answer is yes, the relationship is healthier because both sides know what evidence is shared.
The final design choice is which standards to create internally before industry standards arrive. Episode IDs, tactile and force fields, safety-envelope versions, worker-consent status, and model-release records can be normalized across cells even if vendors differ. Internal standards do not need to be perfect. They need to be stable enough that a later vendor format can be mapped into the plant's process truth.
The ownership review should end with one uncomfortable question: what would the plant know if the vendor system went offline for a week? If operators, quality, maintenance, and IT could still explain recent failures and protect the line, ownership is real. If the answer sits only in a vendor dashboard, the plant has bought capability without retaining memory.
Manufacturing Cell Checkpoint
| Stage | Artifact the manufacturer must own | What to require from vendors | Internal owner |
|---|---|---|---|
| Before PoC | Task taxonomy, defect codebook, episode schema | Proposed export fields and data-use boundaries | Process engineering, quality, IT/security |
| During PoC | Frozen replay set, operator-correction log, QA trace | Before/after replay results and failure explanation | Line owner, robot engineer |
| Limited deployment | Safety envelope, rollback threshold, maintenance-linked logs | Release notes, rollback path, support SLA | Safety, quality, maintenance |
| Scale-up | Vendor exit test, migration format, consent/retention audit | Model-independent export and documentation | Operations lead, legal, procurement |
The point is that ownership is distributed. Process engineering, quality, safety, IT, procurement, legal, and production all hold part of the flywheel. Large-data manipulation often fails because these responsibilities are separated, not only because a model is weak.
Open Questions and Failure Modes
First, the manufacturer can hand the full data flywheel to a vendor and lose the explanation for process improvement. Second, public dataset scale can be mistaken for the QA-linked replay needed in a plant. Third, worker video and process IP can be used for learning before consent and retention rules are ready. Fourth, model updates and hardware maintenance can live in separate logs, making performance shifts hard to diagnose.
The open question is the lack of industrial standards. Robot episode export formats, tactile and force log schemas, update audit trails, and worker-consent models are not yet standardized across manufacturing. Until they are, leading manufacturers need their own internal standards. Later vendor or industry standards will be easier to adopt if process truth is already organized.
Closing
Large-data driven manipulation is ultimately an ownership problem. A manufacturer can buy a model, replace a hand, or switch vendors. If it loses task definitions, failure taxonomies, replay sets, QA traces, safety cases, and data governance, it has purchased automation without owning a learning process. The manufacturers that keep those assets can turn foundation models and production flywheels into durable competitive advantage.
What to Learn Next
The next layer of work is not another model leaderboard. It is the industrial standardization work around episode export, force and tactile log schemas, vendor rollback rights, and worker-consent preserving correction data. Those questions decide whether large-data driven manipulation becomes a repeatable manufacturing capability rather than a strong vendor demonstration.
References
- Open X-Embodiment Collaboration et al. (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv.
- Khazatsky, Alexander (2024). DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset. arXiv.
- AgiBot-World-Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv.
- Chi, Cheng (2024). Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. arXiv.
- Choi, Hojung (2026). In-the-Wild Compliant Manipulation with UMI-FT. arXiv.
- Black, Kevin (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. arXiv.
- Physical Intelligence (2025). OpenPI: Open Source Robot Policy Stack. GitHub.
- Covariant (2024). Covariant Introduces RFM-1 to Give Robots the Human-like Ability to Reason. First-party-distributed company release.
- Dexterity (2025). Dexterity Foresight: AI Platform for Industrial Robot Workcells. Company product page.
- Chef Robotics (2025). ChefOS: AI Robotics Platform for Food Manufacturing. Company product page.
- Figure AI (2026). Figure 03 + Helix 02: General-Purpose Humanoid System. Company product page.
- Generalist AI (2025). GEN-0 Robot Foundation Model. Company page.
- Skild AI (2024). General-Purpose Robot Brain. Company page.
- NVIDIA (2026). NVIDIA Agent Toolkit Expands With New Omniverse Libraries. First-party release.
- Gao, Sicong et al. (2026). NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics. arXiv survey.
- Toyota Research Institute (2023). Toyota Research Institute Unveils Breakthrough in Teaching Robots New Behaviors. Company technical post.
- Li, Xingyu (2024). Evaluating Real-World Robot Manipulation Policies in Simulation. arXiv.
- Mandlekar, Ajay (2021). What Matters in Learning from Offline Human Demonstrations for Robot Manipulation. arXiv.
- Zhao, Tony Z. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv.
- Chi, Cheng (2023). Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. arXiv.
- Yu, Jiawen (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. arXiv.
- Huang, Yuhang (2025). Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization. arXiv.
- Shaw, Kenneth (2023). LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning. arXiv.
- Lambeta, Mike (2024). Digitizing Touch with an Artificial Multimodal Fingertip. arXiv.
- Liu, Huihan et al. (2023). Model-Based Runtime Monitoring with Interactive Imitation Learning. arXiv.
- Dulac-Arnold, Gabriel et al. (2021). Challenges of Real-World Reinforcement Learning: Definitions, Benchmarks and Analysis. Machine Learning.
- Desai, Ankush et al. (2019). SOTER: A Runtime Assurance Framework for Programming Safe Robotics Systems. DSN.
- Dan, Prithwish et al. (2025). X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real. arXiv.
- Lum, Tyler Ga Wei et al. (2025). Crossing the Human-Robot Embodiment Gap with Sim-to-Real RL using One Human Demonstration. arXiv.
- Nguyen, Duc Huy et al. (2024). TacEx: GelSight Tactile Simulation in Isaac Sim -- Combining Soft-Body and Visuotactile Simulators. arXiv.
- Xiao, Wenli et al. (2026). ENPIRE: Agentic Robot Policy Self-Improvement in the Real World. arXiv. #69
- Ranawaka, Nadun et al. (2026). SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation. arXiv preprint.