Kling came from a short-video company, not a lab
Kuaishou, TikTok's main rival in China, released a video model that rivals Sora's samples, and ordinary users in China can already try it. Video models get built by whoever has the video.
Last week Kuaishou released Kling, a text-to-video model, and opened a waitlist inside its creative app in China. The samples are the first I’ve seen that compare to Sora’s. A Chinese man eating noodles with chopsticks, with convincing hand movements and noodles that behave like noodles. A cat driving a car. Clips up to two minutes at 1080p and 30 frames per second, according to Kuaishou. Several people I know in China got access within days and have been sending me clips on WeChat.
Sora was announced in February and, four months later, is still available only to a small group of artists and testers. Kling, from a company most people outside China haven’t heard of, is in the hands of regular users.
What Kuaishou has said about the architecture sounds similar to Sora’s: a diffusion transformer operating on video compressed by a 3D VAE, which compresses across space and time together, and variable aspect ratios and durations. There’s no reason to think the architecture is a secret advantage. What Kuaishou has is data and distribution.
Kuaishou runs one of the two biggest short video platforms in China, with hundreds of millions of daily users uploading enormous amounts of video every day, most of it real people doing real things: eating, cooking, working, dancing, fixing cars, farming. For a video model that has to learn how people move and how everyday physics looks, that’s exactly the data you’d want. And Kuaishou has the creators to put the model in front of immediately.
I think this is a pattern. Video models will be built by companies that already own large amounts of video and large audiences of creators. The foundation labs have the research talent and the compute. But the long-term advantages in video are the data to train on, the users to learn from, and a product where the output is actually used. ByteDance obviously has all three too, and I’d be surprised if it doesn’t release something soon.
The other thing I’ve started paying attention to is cost. From what I hear from friends there, generation in China is much cheaper for users than anything planned in the US, partly because of competition among many Chinese companies building video models, and partly because they’re treating it as a creator tool to drive engagement, not a premium subscription. If that holds, the economics of AI video production could be set in China first. That matters a lot for anyone thinking about AI-made content at scale, like short drama series, where the cost per minute of finished video decides what’s viable.