The decade cameras learned to see
Looking back at 2010 to 2019 from the end of it: the moments that mattered to me, what I got right and wrong on this blog, and one bet for the 2020s.
At the start of this decade I was an undergraduate in Wuhan studying gearboxes, and a computer couldn’t reliably tell a cat from a dog in a photo. Now phones recognize their owners’ faces in the dark, translate speech, and generate faces of people who don’t exist. I’ve been writing here since 2014, so I’ll use the last week of the year to look back.
The moments that changed things, as I saw them:
AlexNet in 2012, when a neural network halved the error on ImageNet and every hand-engineered vision system started to look old. It ended the approach I did my master’s in, and I’m fine with that now.
The DQN Atari paper, 2013 and 2015, the first system I saw learn by doing.
Move 37 in 2016, and then AlphaGo Zero in 2017 learning Go from nothing and beating every version trained on human games.
The Transformer in 2017, which I guessed would spread beyond translation, and then GPT and BERT in 2018 turning it into the default for language.
Face ID in 2017, when 3D face capture went from a research problem to something hundreds of millions of people use every day without thinking.
Musical.ly becoming TikTok, which showed that a recommendation engine can be the product.
What I got right here: pressure-sensitive phones, pretraining taking over language, Transformers spreading, TikTok’s growth. What I got wrong: I said Magic Leap would ship something impressive to consumers within two years. It shipped a developer headset in 2018 that disappointed almost everyone. I said conversational interfaces were at least a decade away in 2016, and I’m starting to suspect I’ll be wrong about that one too, in the other direction, given what GPT-2 can do. On self-driving, I said in 2016 that fully driverless service by 2021 was too optimistic outside small, well-mapped areas, and that still looks right.
What didn’t happen this decade, and what I think the next one is about: machines that understand the physical world. Everything above is about perception, language, games or recommendation. Nothing we’ve built can pick up an unfamiliar object as well as a toddler, or predict what will happen if you push a glass off a table in a way that generalizes to a glass it’s never seen. The systems learned from images, text and game scores. None of them had a body.
So my bet for the 2020s: the breakthroughs that matter most will come from systems that learn by interacting with the world, in simulation first and then in reality, with bodies and hands and more than one sense, driven by something like curiosity more than by a score. The pieces exist now in separate papers: world models, curiosity, domain randomization, dexterous hands. Someone will put them together.