Freiburg-based Black Forest Labs on September 23 released FLUX 3 Action, an open-weight 7-billion-parameter world-action model that turns camera frames and proprioception into millisecond-latency robot control. The model hits 42.92% on the RoboLab benchmark and 93.3% single-item success on a Franka Emika arm.
From Pictures To Actions
Black Forest Labs made its name with the FLUX text-to-image family. FLUX 3 Action is the lab's first foray into vision-language-action (VLA) modeling — the class of models increasingly baked into humanoid and manipulator stacks. It shares the tokenizer with FLUX 3 image but adds a proprioception encoder and an action head trained on ~180 million teleoperation trajectories.
Costs Come Down
Black Forest Labs reports rollout costs of $0.087 per pass on a single H200 and roughly $0.0018 per successful sample once parallelized across eight GPUs. Those numbers put open-weight physical-AI training within reach of research labs and mid-size robotics startups for the first time — a segment that had been priced out by proprietary VLA APIs.
How It Slots Into Physical AI
FLUX 3 Action lands alongside a wave of physical-AI models from NVIDIA Isaac GR00T, Figure Helix, and Boston Dynamics-DeepMind. It also arrives as Qualcomm buys PickNik Robotics to own its manipulation stack and as Figure AI's Helix 2.5 begins home trials.
Reporting based on coverage from Black Forest Labs, Hugging Face and industry announcements on September 23, 2026.
