Vinson·Li

Essay No. 87

You can't skim a song

Why music recommendation is a different problem from video or news recommendation: sequences, repetition, passive listening, and the difference between a skip and a mood.


Two months into working on music, the thing I keep explaining to friends who work on other recommenders is that music breaks most of the assumptions they’re used to. This is all from general principles and public research, not anything specific to my work, but I think it’s worth writing down.

You can’t skim a song. A news article can be judged in a few seconds of scrolling. A video’s thumbnail and first few seconds tell you a lot. A song takes time, and its best part might be a minute in. So the cost of a bad recommendation is higher. The listener spends thirty seconds, not three, finding out it’s wrong, and often they don’t even notice until later, because they weren’t paying attention.

Repetition is the point. In most recommenders, recommending something the user already consumed is a failure. In music, most listening is to songs people already know. A good music recommender has to balance familiarity and discovery, and the right balance depends on the person, the time of day and what they’re doing. Someone at the gym wants their favorites. Someone on a Sunday morning might want something new. The same person might be both.

Signals are ambiguous. If someone plays a playlist while working and never touches the app for two hours, did they love every song, or were they not listening? A skip is a strong negative, but it might mean “not now” rather than “never.” A replay is strong, but people replay songs for reasons that have nothing to do with the recommender. Music has fewer clean signals per hour than short video, where every swipe is a clear decision.

Order matters. A good playlist is a sequence, with transitions, energy that builds and falls, a sense of arc. Ranking songs independently by predicted score and playing them in that order gives you a pile, not a playlist. That makes me think the right model for listening is closer to a language model than a ranker: a session is a sequence, each song depends on the ones before it, and what you want is a good next song given where the sequence has been.

Context is half the answer. Morning commute, workout, dinner, falling asleep. The best signal for what someone wants to hear might be when and where they are, more than their long-term taste.

None of this is news to people who’ve worked in music for years. But coming from faces and events, I’m finding it a fascinating problem. It’s also a humbling one. Music is where people’s identities live, and getting it wrong feels more personal to a listener than a bad video suggestion does.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…