The model plans the shots now
ByteDance's Seedance 2.0 generates multi-shot sequences with references for characters, props and sound, and triggered cease-and-desist letters within a day. What's left for the editor, and for rights holders.
ByteDance released Seedance 2.0 last Thursday, and within a day the internet was full of clips people had made with it, including a now-infamous fight between two very recognizable Hollywood stars. By Friday Disney had sent ByteDance a cease-and-desist letter, and by the weekend the Motion Picture Association and SAG-AFTRA had condemned it publicly. Paramount has reportedly followed.
I work on generative video at YouTube. This is my reaction to the public release, both to what the tool can do and to the disputes around it.
The first is capability. Seedance 2.0 produces clips up to 15 seconds with synchronized audio, and it’s built around references. You can give it many images, video clips and audio clips, for a character’s face, a costume, a location, a camera move or a voice, and it uses them together in one generation. It also plans multiple shots inside one clip: a wide, then a close-up, then a reverse, with the character and setting consistent across the cuts. You can edit specific moments without regenerating the whole thing.
In 2023 I wrote that generated video didn’t cut together because nothing held the scene: faces changed between shots, eyelines didn’t match, screen direction was random. Seedance 2.0 handles a lot of that inside one generation. The model is doing some of the job of a director and editor, deciding how to cover a moment with several shots and keeping them coherent. For short-form storytelling, including the microdramas I wrote about last summer, that’s a big change. A 15-second beat with three shots and dialogue used to take a team hours of generation and editing. Now it’s one prompt with good references.
There’s still plenty for the person to decide. Choosing the references, which is casting and production design. Deciding what the beat is and how it connects to the next one, which is writing and editing. Judging which of several takes has the right performance, which is directing. The model plans the shots within a moment. Someone still has to plan the moments.
The second thing is rights, and it’s the more urgent one. The speed and quality with which people recreated real actors and studio characters show that the model learned them very well, and the industry’s reaction shows nobody agreed to that. I’ve written about consent debt since the face datasets in 2019, and about voices with Fake Drake in 2023. This is the same debt at the scale of the entire film industry, and the actors whose likenesses are the most valuable are the ones most exposed.
I’d expect this to push the whole field toward licensing and consent systems faster than any court case would. Tools will have to verify who’s in the references, block likenesses without permission, and share value with people whose faces and voices are used. The technology for multi-shot storytelling is essentially here. Whether it’s allowed to be used broadly depends on getting that part right.