Vinson·Li

Index

On diffusion

  1. A neural net runs Doom

    Google Research's GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.

    2 min
  2. Control beats prompts

    ControlNet lets you steer Stable Diffusion with a pose skeleton, a depth map or an edge sketch. Creators want to set the structure directly, and this gives them a way to.

    2 min
  3. Text to video is next

    Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.

    2 min
  4. Stable Diffusion runs on my own computer

    Stability AI released the weights of a text-to-image model anyone can run on a consumer GPU. Open models change who gets to build, and what gets built.

    2 min
  5. Images are solved-ish. Video is where physics lives

    DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.

    2 min
  6. Diffusion is going to eat GANs

    Two papers this week, GLIDE and latent diffusion, make text-to-image generation with diffusion models look practical. What denoising actually learns, and why it beats the adversarial game.

    2 min