Vinson·Li

Index

On representation

  1. Gaussian splatting killed my mesh nostalgia

    3D Gaussian Splatting represents a scene as millions of fuzzy, colored blobs and renders it in real time at NeRF quality. It does this without a neural network.

    2 min
  2. Predicting in representation space

    Meta's I-JEPA is the first concrete result from LeCun's world model agenda. It learns image representations by predicting hidden regions in latent space, with no augmentations and no pixel reconstruction.

    2 min
  3. Recommendation as next-token prediction

    A new paper turns every item into a short code of semantic tokens, then has a Transformer generate the code of what you'll want next. The item vocabulary finally describes what things are.

    2 min
  4. Text to music is a representation problem

    Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.

    2 min
  5. NeRF went from hours to seconds

    Nvidia's Instant NGP trains a neural radiance field in seconds with a multiresolution hash table. The representation mattered more than the network.

    2 min
  6. Pictures and words in the same space

    OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.

    2 min
  7. A whole scene stored inside a network

    NeRF represents a 3D scene as a small neural network you can render from any viewpoint. No mesh, no triangles. My old thesis problem, answered from a different direction.

    2 min
  8. BERT reads both directions at once

    Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.

    2 min