On reinforcement-learning
- Search plus verification
AlphaProof and AlphaGeometry 2 together solved four of six problems from this year's Math Olympiad, reaching silver-medal level. What makes a checkable answer so useful for training, and what happens when there isn't one.
2 min reads likes comments - MuZero learns the rules it isn't given
DeepMind's MuZero plays Go, chess, shogi and Atari at top level without being told the rules. It plans inside a model it learned, and the model only predicts what matters.
2 min reads likes comments - The noisy TV problem
A curious agent gets hypnotized by random static. OpenAI's Random Network Distillation fixes it and beats humans at Montezuma's Revenge. What curiosity should actually reward.
2 min reads likes comments - A robot hand with a hundred years of practice
OpenAI's Dactyl learned to rotate a block with a human-like robot hand, entirely in simulation. Why hands, and why randomizing the simulator is the clever part.
2 min reads likes comments - An agent that dreams its own racetrack
Ha and Schmidhuber's World Models compresses what an agent sees, learns to predict what happens next, and trains a tiny controller inside its own dream. The most important paper I've read this year.
2 min reads likes comments - AlphaGo Zero threw away human games and got better
Starting from random play, with no human data, it beat the version that beat Lee Sedol 100 games to 0 after three days. Human knowledge was a ceiling.
2 min reads likes comments - Nobody taught it to walk
DeepMind's locomotion agents learned to run, jump and duck from a reward for moving forward and a varied course. The flailing arms are the point.
2 min reads likes comments - Curiosity is a prediction error
A Berkeley paper gives agents an internal reward for being surprised by their own predictions. It learns Mario with no score at all, and it answers a question I asked two years ago.
2 min reads likes comments - Every lab is building a playground
DeepMind open-sourced its 3D lab and OpenAI released Universe in the same week. Environments are becoming the new datasets.
2 min reads likes comments - Move 37
AlphaGo beat Lee Sedol four games to one. The move everyone is talking about, how the system found it, and why self-play is the part that matters.
2 min reads likes comments - 49 Atari games from pixels, and zero points in Montezuma's Revenge
DeepMind's DQN paper is in Nature. What the network actually learns, and the game where it learns nothing at all.
2 min reads likes comments