Vinson·Li

Essay No. 89

Listening sessions are sentences

Treat each song as a token and each listening session as a sentence, and recommendation starts to look like language modeling. What that framing gets right, and what it misses.


In July I wrote that a good playlist is a sequence, and that the right model for listening might be closer to a language model than a ranker. I’ve been reading the public research on that idea since, and I want to lay out how it works, because I think it’s one of the more important shifts in recommendation.

The traditional setup, roughly, is a ranker. For a user and a candidate item, predict a score, like the chance they’ll click or listen through. Score many candidates, sort, show the top ones. The user’s history goes in as features, often summarized: the genres they like, their top artists, an average of the embeddings of things they’ve played. The order of their history mostly gets lost in that summary.

The sequential approach keeps the order. Give every item an ID and a learned embedding, the same way a language model gives every word one. A user’s history becomes a sequence of item tokens: the last fifty songs they played, in order. Train a model to predict the next item from the ones before it, exactly like predicting the next word in a sentence. SASRec in 2018 did this with a Transformer decoder, GPT-style. BERT4Rec in 2019 did it BERT-style: hide some items in the sequence and predict them from both sides.

Why this helps is the same reason it helps in language. The meaning of a song in your session depends on context. A quiet acoustic track after a run of ambient music is part of a focus session. The same track after a run of breakup songs means something else. Attention lets the model weigh which earlier items matter for the next one, so it can pick up short-term intent, like “this person is winding down for sleep,” along with long-term taste.

It’s also important what this isn’t. People hear “Transformer” and think of GPT writing essays. These models don’t generate text or know anything about language. Their vocabulary is the item catalog, possibly millions of songs, and their output is a probability distribution over which item comes next. It’s the same architecture solving a completely different problem, which is a nice example of how general attention turned out to be.

It has real problems. The vocabulary is huge and keeps growing, since new songs arrive every day and a new song has no learned embedding. That’s the cold start problem I wrote about with DLRM, now inside a sequence model. And predicting “the next item” can make a recommender too literal: it learns what people usually played next, which reinforces what’s already popular.

The way out, I think, is to connect the item tokens to content. If a song’s embedding came from its audio, lyrics and metadata, not only from its ID, a new release would start in a sensible place. Even better would be item “words” that are built from content and shared across items, so the model’s vocabulary doesn’t need one entry per song. Then recommendation would really look like language modeling, with a vocabulary that describes what things are instead of just naming them.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…