Vinson·Li

Index

On computer-vision

  1. Pictures and words in the same space

    OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.

    2 min
  2. An image is worth 16x16 words

    A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.

    2 min
  3. The black hole photo was a reconstruction problem

    The Event Horizon Telescope didn't take a picture in the ordinary sense. It filled in a mostly empty measurement with priors, and was careful about which ones.

    2 min
  4. 152 layers, and the trick is learning nothing

    Microsoft Research's residual networks won ImageNet with a network eight times deeper than last year's. The idea behind it is almost too simple.

    2 min
  5. DeepDream sees dogs everywhere

    Running a network in reverse to see what it learned. Why everything turns into dogs, and what that says about training data.

    2 min
  6. HoloLens maps your room before it draws on it

    Microsoft's headset is mostly a sensing problem with a display attached. Why AR needs a model of the space before it can put anything in it.

    2 min
  7. A computer captioned a photo. It didn't see the photo

    Google and Stanford both have networks that write sentences about images. How they work, and what they're actually learning.

    2 min
  8. A video is not a stack of photos

    Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.

    3 min