On deep-learning
- An image is worth 16x16 words
A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.
2 min reads likes comments - A whole scene stored inside a network
NeRF represents a 3D scene as a small neural network you can render from any viewpoint. No mesh, no triangles. My old thesis problem, answered from a different direction.
2 min reads likes comments - The embedding table is the model
Facebook open-sourced DLRM, its deep learning recommendation model. It shows what big recommenders actually look like: mostly memory, with a small network on top.
2 min reads likes comments - Read everything first, specialize later
OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.
2 min reads likes comments - Faces help you hear
Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.
2 min reads likes comments - Can a network have taste?
Google's NIMA predicts how people would rate a photo, as a distribution, not a single score. Modeling disagreement turns out to be the useful part.
2 min reads likes comments - No recurrence, no convolution
A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.
2 min reads likes comments - Google Translate got better overnight, and Chinese speakers noticed first
Neural machine translation replaced phrase tables, starting with Chinese to English. Notes from someone who reads both, and why the zero-shot result is the bigger story.
2 min reads likes comments - 152 layers, and the trick is learning nothing
Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.
2 min reads likes comments - Van Gogh is a Gram matrix
A new paper separates the content of an image from its style using a network trained for classification. How it works, and what it suggests about taste.
2 min reads likes comments - DeepDream sees dogs everywhere
Running a network in reverse to see what it learned. Why everything turns into dogs, and what that says about training data.
2 min reads likes comments - 49 Atari games from pixels, and zero points in Montezuma's Revenge
DeepMind's DQN paper is in Nature. What the network actually learns, and the game where it learns nothing at all.
2 min reads likes comments - Two networks arguing
The GAN paper from NIPS this year: a generator, a discriminator, and a learned idea of what counts as real.
2 min reads likes comments - A computer captioned a photo. It didn't see the photo
Google and Stanford both have networks that write sentences about images. How they work, and what they're actually learning.
2 min reads likes comments - A video is not a stack of photos
Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.
3 min reads likes comments - Graduated. My thesis in one paragraph, and how deep learning will eat it
What my master's thesis on 3D face meshes does, which part of it I think neural networks will replace soon, and which part I think they won't.
2 min reads likes comments