Palo Alto-based Generalist AI has unveiled GEN-1.5, an embodied foundation model that the company says can learn new short-horizon manipulation tasks from as little as three to 12 seconds of demonstration data with no gradient updates — a step that some observers are already calling a “GPT-3 moment for physical AI.”
Physical Prompting In A 30-Second Context Window
Announced on the Generalist blog on August 19, GEN-1.5 is a large multimodal model that ingests video, sensor and proprioceptive inputs across a 30-second context window and produces 100 Hz action trajectories. Generalist calls the technique physical prompting: sensorimotor demonstrations recorded from human hands or robot rollouts are placed in the model’s context, and the model executes the inferred task immediately. Across 10 diverse manipulation tasks — opening jars, handling zippers, retrieving money from wallets — the model averaged 59% success (±10%) with one-shot in-context prompting and 83% (±9%) after 10 gradient steps on five minutes of data per task.

Compositional, Sim-To-Real And Human-To-Robot
The company reports that GEN-1.5 exhibits capabilities the team did not explicitly train for. Two independent physical prompts placed in context can be composed into a single longer-horizon behavior, with the model producing its own bridging motions. Prompts recorded entirely in simulation transfer to the real robot despite zero simulation data in pretraining. And in some cases, a human demonstrating a task with their own hands in view of the robot’s cameras is enough for the robot to reproduce it immediately.
Emerging From Eight Months Of Continuous Pretraining
GEN-1.5’s pretraining has been running continuously for more than eight months on Generalist’s in-house data engine, which captures activities from homes, warehouses and factories. The company said adaptation to new tasks has grown steadily more data- and compute-efficient across GEN-0 and GEN-1, and that GEN-1.5 can now be fine-tuned in one to 10 gradient steps on one to five minutes of data. That level of fast adaptation was previously considered out of reach for robotics.
The release lands in an increasingly crowded physical-AI market. For related coverage, see Acorn Robot’s Natus AGE-0 tactile model, Veeda AI’s $90M seed for world models and Beijing Humanoid’s Tiangong 2.0 with the WU world model.
Reporting based on coverage from Generalist AI, The Decoder, TechTimes and Humanoids Daily.
