Thore Graepel, one of the researchers who taught Google DeepMind's AlphaGo to beat the world's best Go player, has quietly quit DeepMind to build a new AI reasoning startup called Metis Reasoning — and is already courting investors for a first round in the "tens of millions" of dollars.
Who Is Graepel
Graepel joined DeepMind in 2015, roughly a year before AlphaGo defeated Lee Sedol in Seoul. He is credited as a key architect of AlphaGo, AlphaZero and MuZero — a lineage of systems that used deep reinforcement learning plus Monte Carlo tree search to solve problems no one had cracked before. He left DeepMind in 2021 for biology-focused Altos Labs, returned to the lab in mid-2025, and, according to Bloomberg, departed again in the summer to launch Metis Reasoning.

The Pitch: Search Beats Scale
Metis is chasing "structured search methods and probabilistic planning" — the AlphaZero playbook — rather than pouring more tokens into transformer training. Graepel argues the winners of the next AI cycle will be systems that can plan under uncertainty, an area where today's LLMs still hallucinate their way through robotics, protein design or logistics optimization. Applications named include science, engineering and robotics.
Round Size
Sources tell Bloomberg the London-based startup is targeting an initial round in the tens of millions with valuations rising into the hundreds of millions in the next twelve months. Investors have been circling ex-DeepMind founders including Anthropic-adjacent talent and Physical Intelligence.
Why It Matters For Robotics
Reasoning-native AI is exactly the missing layer that most robotics developers — 1X or Black Forest Labs among them — still bolt onto their vision-language-action stacks. If Graepel's new company converts AlphaGo's planning tricks into a general reasoning API, the humanoid and autonomous-vehicle industries will be first in the queue.
Reporting based on coverage from Bloomberg and Crypto Briefing.