On computer-vision
- Pictures and words in the same space
OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.
2 min reads likes comments - An image is worth 16x16 words
A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.
2 min reads likes comments - The black hole photo was a reconstruction problem
The Event Horizon Telescope didn't take a picture in the ordinary sense. It filled in a mostly empty measurement with priors, and was careful about which ones.
2 min reads likes comments - 152 layers, and the trick is learning nothing
Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.
2 min reads likes comments - DeepDream sees dogs everywhere
Running a network in reverse to see what it learned. Why everything turns into dogs, and what that says about training data.
2 min reads likes comments - HoloLens maps your room before it draws on it
Microsoft's headset is mostly a sensing problem with a display attached. Why AR needs a model of the space before it can put anything in it.
2 min reads likes comments - A computer captioned a photo. It didn't see the photo
Google and Stanford both have networks that write sentences about images. How they work, and what they're actually learning.
2 min reads likes comments - A video is not a stack of photos
Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.
3 min reads likes comments