On generative-models
- Change everything except the motion
At I/O this week, YouTube brought Gemini Omni to Shorts Remix: take a Short, change the characters, setting or style, and keep the motion and story. Why remix is the right first experience for video-to-video.
2 min reads likes comments - Sora is shutting down. Generation isn't a product
OpenAI closed the Sora app this week, seven months after launching it as a social network, and the API ends in September. What went wrong, and what it says about where generative video belongs.
2 min reads likes comments - The model plans the shots now
ByteDance's Seedance 2.0 generates multi-shot sequences with references for characters, props and sound, and triggered cease-and-desist letters within a day. What's left for the editor, and for rights holders.
2 min reads likes comments - Sora built a social network
OpenAI launched Sora 2 as a TikTok-style app where every video is generated. The model is impressive. I'm less sure a feed of generations gives people a reason to keep watching.
2 min reads likes comments - Consistent characters are the unlock
Google's new image model, known everywhere as Nano Banana, keeps a person looking like themselves across edits and scenes. For AI storytelling, identity preservation matters more than image quality.
2 min reads likes comments - Sixty episodes, one minute each
Microdramas are one of the fastest-growing forms of entertainment, and the format exists because of production cost. What happens when AI changes the cost?
2 min reads likes comments - Veo 3 has sound, and that changes the medium
Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.
2 min reads likes comments - What a minute of AI video costs in Beijing vs. San Francisco
Real per-second prices for Veo 2, Kling 2.0 and Jimeng this month, and why the retake rate matters more than the sticker price.
2 min reads likes comments - Style is free now. Taste isn't
GPT-4o's image generation turned the internet into Studio Ghibli for a week. When every style is one prompt away, the scarce thing is judgment.
2 min reads likes comments - Genie 2 and World Labs in the same week
DeepMind's Genie 2 turns one image into a playable 3D world, and Fei-Fei Li's World Labs turns one image into a 3D scene you can walk through. Worlds are the next medium after text, images and video.
2 min reads likes comments - Kling came from a short-video company, not a lab
Kuaishou, TikTok's main rival in China, released a video model that rivals Sora's samples, and ordinary users in China can already try it. Video models get built by whoever has the video.
2 min reads likes comments - Music generation's GPT-3 moment
Suno v3 and Udio generate full songs with vocals and lyrics from a sentence, and some of them are good. What that means for artists, platforms and listeners.
2 min reads likes comments - "World simulator" is doing a lot of work
OpenAI's Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that's tempting and why it isn't true yet.
2 min reads likes comments - Four-second clips can't make a movie
Pika 1.0, Stable Video Diffusion and Runway's latest make beautiful short clips. What's missing is the grammar of film: continuity, screen direction and eyelines.
2 min reads likes comments - The model rewrites your prompt
DALL·E 3 in ChatGPT doesn't use what you typed. It has a language model write a long, detailed prompt for you. Creative tools are turning into agents that interpret intent.
2 min reads likes comments - Fake Drake
An AI-generated song imitating Drake and The Weeknd got millions of plays before it was pulled. Voice is identity, and the industry needs consent and attribution systems, fast.
2 min reads likes comments - Control beats prompts
ControlNet lets you steer Stable Diffusion with a pose skeleton, a depth map or an edge sketch. Creators want to set the structure directly, and this gives them a way to.
2 min reads likes comments - Text to music is a representation problem
Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.
2 min reads likes comments - Text to video is next
Meta's Make-A-Video generates short clips from a sentence. They're five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.
2 min reads likes comments - Stable Diffusion runs on my own computer
Stability AI released the weights of a text-to-image model anyone can run on a consumer GPU. Open models change who gets to build, and what gets built.
2 min reads likes comments - Images are solved-ish. Video is where physics lives
DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.
2 min reads likes comments - Diffusion is going to eat GANs
Two papers this week, GLIDE and latent diffusion, make text-to-image generation with diffusion models look practical. What denoising actually learns, and why it beats the adversarial game.
2 min reads likes comments - "Too dangerous to release"
OpenAI trained a much bigger language model and is holding back the full version. The unicorn story is impressive. What's actually dangerous is cheap, plausible text at scale.
2 min reads likes comments - None of these people exist
Nvidia's StyleGAN generates photographic faces with control over pose, identity and freckles. A face company's view of faces becoming free.
2 min reads likes comments - An agent that dreams its own racetrack
Ha and Schmidhuber's World Models compresses what an agent sees, learns to predict what happens next, and trains a tiny controller inside its own dream. The most important paper I've read this year.
2 min reads likes comments - Face swaps on Reddit: I built the harmless version years ago
Someone on Reddit is putting celebrities' faces into porn with a home GPU and open-source tools. How it works, and why consent has to be designed in now.
2 min reads likes comments - 16,000 samples a second
DeepMind's WaveNet generates raw audio one sample at a time, and its speech sounds far more human than anything before. How, and why it's so slow.
2 min reads likes comments - Van Gogh is a Gram matrix
A new paper separates the content of an image from its style using a network trained for classification. How it works, and what it suggests about taste.
2 min reads likes comments - Two networks arguing
The GAN paper from NIPS this year: a generator, a discriminator, and a learned idea of what counts as real.
2 min reads likes comments