Vinson·Li

Essay No. 145

From embeddings to IDs to what you watch next

Generative recommenders are going into production. How a video becomes an embedding, then a semantic ID, and how a Transformer uses those IDs to choose what you'll watch next.


In 2021 I wrote that listening sessions are sentences, and in 2023 that recommendation was becoming next-token prediction, after the TIGER paper introduced semantic IDs. Since then the idea has gone from research to production at several companies. Kuaishou has published on OneRec, an end-to-end generative recommender serving real traffic, and Meta has written about its generative recommender work. I get asked often enough how these systems actually work, mechanically, that I want to write it down once. This is from public papers and general principles. It doesn’t describe any specific system I work on.

Step one: understand the content. Each video is run through a multimodal model that looks at frames, audio, speech transcript, title and description, and produces a content embedding, a vector of a few hundred to a few thousand numbers. The key point is that this is about the video itself, not about who watched it. Two videos about fixing the same bike problem land close together even if they were uploaded an hour ago and nobody has watched either. How you sample the video matters. A few thumbnails miss what happens over time, which I wrote about in 2014 with optical flow, and it’s still true.

Step two: turn embeddings into tokens. A vector is continuous, and a Transformer that generates sequences needs a discrete vocabulary. So you quantize. Residual quantization is the usual approach: the first codebook assigns the embedding to its nearest of, say, a few thousand centroids, and that’s the first token. Subtract the centroid and quantize the residual with a second codebook to get the second token. Repeat three or four times. Each video now has a short code, its semantic ID, where the first token is coarse (“cooking”) and later ones are finer (“Sichuan, home kitchen, fast-paced, one presenter”). Similar videos share prefixes. When two videos collide on the full code, you add a disambiguation token.

Step three: train on sequences. A user’s history becomes a long sequence of semantic ID tokens, plus tokens for context and actions: watched to the end, skipped at two seconds, liked, shared. The model, a Transformer, is trained to predict the next video’s ID tokens given everything before. Where earlier systems had a separate retrieval stage and ranking stage, these models increasingly do both: they generate candidates directly, and they’re trained or fine-tuned with preference signals so what they generate is what the user will actually value, which is often different from what they’d click.

Step four: from IDs back to videos. At serving time, the model generates the next semantic ID one token at a time, usually with beam search so you get many candidate codes. Each code maps to a small set of actual videos in an index. Some generated codes won’t exist. You constrain decoding to valid prefixes, with a trie of real IDs, so the model only generates codes that point somewhere.

Why this is worth the complexity: new videos can be recommended the moment they’re understood, because their semantic IDs come from content, not interaction history. The vocabulary is compact and meaningful, which scales better than billions of per-item embeddings. The model can generate a sequence, the next several videos as a coherent session, instead of scoring items one at a time. And language and items live in the same token space, so “show me something like this but calmer” is a natural input.

There are still problems: codebooks go stale as content shifts, popular items dominate the training sequences, and a generative model can be confidently wrong in very fluent ways. But I’m convinced this is where large-scale recommendation is heading, and it has taken about ten years for the idea of “songs as words” to become how serious systems are built.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…