On transformer
- From embeddings to IDs to what you watch next
Generative recommenders are going into production. How a video becomes an embedding, then a semantic ID, and how a Transformer uses those IDs to choose what you'll watch next.
3 min reads likes comments - Recommendation as next-token prediction
A new paper turns every item into a short code of semantic tokens, then has a Transformer generate the code of what you'll want next. The item vocabulary finally describes what things are.
2 min reads likes comments - One network, 604 tasks
DeepMind's Gato plays Atari, captions images, chats and stacks blocks with a real robot arm, all with the same weights. Turning actions into tokens is the interesting part.
2 min reads likes comments - Listening sessions are sentences
Treat each song as a token and each listening session as a sentence, and recommendation starts to look like language modeling. What that framing gets right, and what it misses.
2 min reads likes comments - One architecture, any input
DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.
2 min reads likes comments - An image is worth 16x16 words
A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.
2 min reads likes comments - BERT reads both directions at once
Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.
2 min reads likes comments - Read everything first, specialize later
OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.
2 min reads likes comments - No recurrence, no convolution
A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.
2 min reads likes comments