Vinson·Li

Essay No. 108

Recommendation as next-token prediction

A new paper turns every item into a short code of semantic tokens, then has a Transformer generate the code of what you'll want next. The item vocabulary finally describes what things are.


In 2021 I wrote that sequential recommenders treat songs as tokens and sessions as sentences, and that the way forward was item “words” built from content and shared across items, so the vocabulary doesn’t need one entry per item. A paper on arXiv this month does something very close to that: “Recommender Systems with Generative Retrieval,” by Rajput and others. They call the system TIGER. The authors are at Google, where I work; this post is based on their public paper.

The standard two-stage recommender works like this. Retrieval uses embeddings: a user vector and item vectors in the same space, and approximate nearest neighbor search to find a few hundred candidates out of millions. Then a heavier ranker scores those. Items are represented by IDs with learned embeddings, and a new item is invisible until it collects interactions.

TIGER replaces retrieval with generation, using what the paper calls Semantic IDs.

First, each item gets a content embedding from a pretrained text encoder, based on its title, description, category and so on. Then those embeddings are quantized into a short tuple of discrete codes with a residual-quantized variational autoencoder (RQ-VAE). The first code picks the nearest of, say, 256 cluster centers. The second code quantizes what’s left over, the residual, with its own codebook. Then a third. So an item becomes something like (17, 203, 45): a coarse category, then finer and finer distinctions. Items with similar content share prefixes, like words sharing roots. An extra code is added only to separate items that collide.

Second, a Transformer encoder-decoder is trained on user histories: the input is the sequence of Semantic IDs of items the user interacted with, and the target is the Semantic ID of the next item, generated one code at a time. At inference, you run beam search over codes and look up which real items those codes correspond to.

What I like about this:

The vocabulary is small and meaningful. Instead of millions of item embeddings, the model has a few codebooks of a few hundred codes each. Every code carries content information, so the model is predicting “what kind of thing, specifically” rather than memorizing IDs.

Cold start gets easier. A brand new item gets a Semantic ID from its content right away. It shares prefixes with similar items that the model already knows how to predict, so it can be recommended before anyone has interacted with it. The paper shows improvements on cold-start items.

It unifies recommendation with the rest of the Transformer world. Once items are sequences of tokens, you can in principle put them in the same sequence as text: “the user said they want something relaxing,” then a history of item codes, then generate the next item. Recommendation and conversation become the same kind of problem.

The open questions are big. The paper is on relatively small public datasets, not billions of items. The quality of recommendations depends entirely on how good the quantization is, and whether content embeddings capture what actually makes people like things. And beam search over codes may generate codes that don’t match any real item. But I think this is the beginning of recommenders that are generative models over a vocabulary of content.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…