Vinson·Li

Essay No. 55

The noisy TV problem

A curious agent gets hypnotized by random static. OpenAI's Random Network Distillation fixes it and beats humans at Montezuma's Revenge. What curiosity should actually reward.


Montezuma’s Revenge keeps showing up on this blog. In 2015 the DQN agent scored zero on it because it never found a reward. Last year I wrote about curiosity as prediction error, rewarding an agent for being surprised by its own forward model. This year OpenAI and Berkeley ran curiosity at large scale, and the story got more interesting.

In August, Burda, Edwards, Pathak and others published “Large-Scale Study of Curiosity-Driven Learning.” They trained agents on 54 environments with no external reward at all, only curiosity. A lot of them learned useful behavior anyway. Agents played long stretches of Atari games and navigated 3D mazes just because new things were surprising.

They also demonstrated the failure everyone suspected. They put a TV in a maze that shows random images when the agent presses a button. A curiosity-driven agent finds the TV and stops exploring. Every channel change is unpredictable, so every button press gives a big intrinsic reward, forever. The agent becomes a couch potato. They call it the noisy TV problem.

It’s a real problem because it’s about the difference between two kinds of unpredictability. Some things are unpredictable because you haven’t learned them yet, like what’s behind a door you haven’t opened. Others are unpredictable because they’re random, like static, dice, or leaves in the wind. Curiosity is only useful for the first kind. The second kind is a trap, because no amount of learning reduces the surprise.

Last month OpenAI’s “Random Network Distillation” offered a neat way around it. Take a randomly initialized network and freeze it. For each observation, it produces some arbitrary output. Train a second network to predict that output. The prediction error is high on observations that are unlike anything you’ve trained on, and it falls as you see more similar observations. It measures novelty, how new this observation is, instead of how hard it is to predict the future from it. Since the target is a deterministic function of the current observation, random transitions don’t produce endless error. With this, their agent explored most of the rooms in Montezuma’s Revenge and beat the average human score, without any demonstrations.

I think the general lesson is that curiosity needs to reward learnable surprise: things where more experience would reduce your uncertainty. That’s also roughly how good students and good engineers work. You don’t want someone fascinated by noise. You want someone drawn to what they don’t understand yet but can.

It’s worth saying how far this still is from a child. These agents explore rooms of a video game. A child explores gravity, faces, language and their own body, all at once, through a body with dozens of senses. But the principle of exploring whatever you can learn from seems like the right principle to scale up.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…