Search plus verification
AlphaProof and AlphaGeometry 2 together solved four of six problems from this year's Math Olympiad, reaching silver-medal level. What makes a checkable answer so useful for training, and what happens when there isn't one.
On Thursday Google DeepMind announced that two of its systems, AlphaProof and AlphaGeometry 2, together solved four of the six problems from this month’s International Mathematical Olympiad, scoring 28 out of 42 points. That’s silver medal level, one point below gold. The solutions were graded by prominent mathematicians using the official rules. Some of the problems took the system up to three days, far beyond the time limit students get. Still, it’s the first time an AI system has reached medal level at the IMO.
AlphaProof is built the way I’d expect DeepMind to build it. It works in Lean, a formal proof language where every step of a proof can be checked mechanically by a computer. A language model, Gemini-based, was used to translate about a million informal math problems into formal Lean statements. Then AlphaProof trained by the AlphaZero method: generate candidate proof steps, search, and use Lean’s checker to verify whether the proof is correct. Correct proofs become training data. The system gets better by practicing on problems it generates, with a perfect judge telling it when it’s right.
It’s AlphaGo Zero again. In 2017 I wrote that the most valuable thing you can build for a problem is a good simulator and a good evaluator, and that with both, self-play will probably beat what humans can teach. Go has the rules as a simulator and the winner as evaluator. Formal math has Lean as both. When you have that, you can generate unlimited experience and learn from it.
This is the part I keep thinking about: which problems have a verifier and which don’t.
Math has one, if you formalize it. Code has a partial one: tests, compilers, whether the program runs. Games have one. A lot of science has a slow, expensive one: the experiment.
Taste doesn’t. Is this photograph beautiful? Is this edit well paced? Is this song good? Is this the right shot to cut to? There’s no Lean for aesthetics. Different people disagree, the same person disagrees with themselves on different days, and the answer depends on context and culture. The AlphaZero recipe doesn’t directly apply, because the “correct” signal doesn’t exist.
Most of the creative applications I care about, music, video, filmmaking, live on that side of the line. The best we have there is learned preference models, trained on which outputs people pick, like NIMA for photos in 2017 or the reward models used in RLHF for chat. Those are noisy, biased toward whoever did the rating, and easy for an optimizer to exploit. A model trained hard against a learned taste model will find the things the taste model likes that people don’t.
I don’t think that’s a dead end. People develop taste too, with no verifier, through exposure, feedback and comparison. But it’s a harder, less settled problem than theorem proving, and I think it’s where a lot of the interesting work in creative AI will be for the next few years.