Reward AI Exits Stealth With OM-1, A Human-Data-Only Robot Foundation Model

Stanford spin-out Reward AI has broken cover with OM-1, a robot manipulation policy trained entirely on human wearable data — no teleoperation, no on-robot pretraining — that transfers zero-shot to tabletop arms, industrial manipulators and bipedal humanoids.

Reward AI Exits Stealth With OM-1, A Human-Data-Only Robot Foundation Model

Stanford University DexCap spin-out Reward AI has come out of stealth with OM-1, a general-purpose robot manipulation policy trained entirely on human demonstration data — no teleoperation, no on-robot pretraining and no internet video. The company published the model, a technical report and a live demo reel on September 14, 2026.

The Trick: Skip The Robot For Training

OM-1 is trained exclusively on data captured through Reward AI's own Omnibody Hand, a seven-degree-of-freedom wearable glove derived from Stanford's DexCap research. The glove fuses electromagnetic tracking, tactile arrays, proximity sensors and in-hand global-shutter cameras. Reward AI clocks its hand-pose accuracy at 9.5 mm mean error against 24.9 mm for visual-inertial-only rigs — the difference between grasping a lug nut cleanly and dropping it on the first try. That data pipeline lets the model absorb a new long-horizon task from under 30 minutes of raw human demonstration.

Reward AI OM-1 policy architecture

Zero-Shot Cross-Embodiment

The value proposition is cross-embodiment transfer. In demos, OM-1 unplugs an Ethernet cable (a task that demands real tactile feedback rather than clever vision), folds laundry bimanually, bartends, packages consumer electronics and reaches into a fridge to retrieve a bottle — with the same model, on different robot bodies, at natural human speed. A separate reinforcement-learning controller runs at a higher clock than the policy, absorbing motor backlash and latency so the policy can plan without micromanaging joint torques.

Where It Sits In The Physical AI Race

Reward AI is positioning itself against two very public rivals. The video-only pretraining camp — Rhoda AI's FutureVision and Dyna Robotics' Dyna-2 — bets the internet has enough manipulation data. The in-context visual-demonstration camp — Skild AI's S1 and Generalist AI's GEN-1.5 — argues you can prompt a policy at inference time. OM-1's answer is that neither substitutes for kinesthetic data collected off a real human hand.

What's Missing

Reward AI has not disclosed funding, revenue, headcount or a headquarters beyond a Palo Alto base. Nor has it named a paying pilot customer, which is the piece the whole physical-AI field is racing to book. As NVIDIA's Les Karpas flagged this week, moving from demo reel to 99.9% uptime across a customer fleet is the industry's shared bottleneck.

Reporting based on coverage from Humanoids Daily, MarkTechPost, Metaverse Post and Reward AI.

Category: Machine Learning

Tags: humanoid robots humanoid AI AI innovation AI perception AI embodiment AI Foundation Models

Related Articles