Zurich-based dexterous manipulation startup mimic robotics has partnered with Germany's Black Forest Labs (BFL) to release FLUX-mimic, a next-generation Video-Action Model (VAM) built on top of BFL's newly announced FLUX 3 multimodal backbone. Announced on July 23, 2026, FLUX-mimic is already being tested on production tasks at Audi and represents mimic's bet that frontier video generative models — not language models — are the fastest path to a general-purpose robot brain.
Control reduces to video prediction
Traditional Vision-Language-Action (VLA) policies bolt an action head onto a language model and are throttled by the scarcity of teleoperated robot data. mimic instead trains its policies on top of a generative video model that has already learned physics and human behavior from billions of frames of human video. In benchmarks on a soft-body kitting task, FLUX-mimic scored 95% success out of the box — no per-task fine-tuning — versus 55% for an adapted pi0.5 baseline and 70% for a heavily post-trained Flow Matching policy. The team says this validates their thesis that “control effectively reduces to video prediction.”
What FLUX-mimic actually does
FLUX 3, unveiled by Black Forest Labs on the same day, is a frontier multimodal generator for image, video and audio. FLUX-mimic keeps the full video backbone trainable and attaches a compact action decoder that attends directly to the backbone's internal representation of predicted future frames, denoising chunks of robot actions without ever rendering pixels. A new task can be fine-tuned with as little as 30 minutes of robot data — down from the 30+ hours conventional pipelines typically require.

Runs on the robot, not in the cloud
Despite the video-model backbone, FLUX-mimic runs locally on a single NVIDIA RTX 5090 through partial video denoising, post-training quantization, Real-Time Chunking that overlaps prediction and execution, and adaptive step-skipping. mimic's own mimic-ipc middleware moves sensor data and actions between hardware with minimal jitter so the policy always acts on a fresh view of the world.
On the Audi line
The model is already meeting the factory floor. Together with Audi's Production Lab, mimic has deployed FLUX-mimic on car-door assembly tasks that traditional automation has never handled — including fitting flexible window-shaft seals (Fensterschachtleiste) and other soft-body manipulation work. “We have seen these robots solve complex soft body manipulation work that would have been simply impossible with conventional robotics,” said Christoph Schneider of Audi Production Lab. Data has been collected across more than 100 real factory use cases through a mix of teleoperation on mimic robots and human motion recorded on the company's wearable capture systems.
Where it slots into the physical-AI race
FLUX-mimic joins a fast-crowding foundation-model shelf for robots — from NVIDIA Cosmos 3 Edge to Genesis AI and Physical Intelligence. Where most rivals extend LLM-derived VLA recipes, the mimic + BFL pairing is a rare full-stack shot at putting a frontier video model on the robot itself.
What's next
FLUX 3 Action — the family of downstream robotics variants including FLUX-mimic — is in early access with selected research and commercial partners. mimic says future versions will accept freeform prompts including motion sketches, goal images and video demonstrations, moving toward robots that can be instructed by showing rather than typing.
Reporting based on mimic robotics, Black Forest Labs and MarkTechPost.
