Vinson·Li

Essay No. 110

Predicting in representation space

Meta's I-JEPA is the first concrete result from LeCun's world model agenda. It learns image representations by predicting hidden regions in latent space, with no augmentations and no pixel reconstruction.


Last July I wrote about Yann LeCun’s position paper and its central idea, the Joint Embedding Predictive Architecture: predict the representation of what’s missing instead of generating the missing pixels. Meta published the first concrete version last month, I-JEPA for images, presented at CVPR. I’ve finally had time to read it properly.

The setup, as I understand it:

Take an image and cut it into patches, as with a Vision Transformer. Pick one large “context” block of patches that the model can see. Pick several smaller “target” blocks elsewhere in the image, which the model can’t see.

A context encoder, a ViT, encodes the visible patches. A target encoder, same architecture, encodes the full image, and you take its outputs at the target block positions. That’s what the model has to predict.

A predictor, a smaller ViT, takes the context representation plus position information saying “the target is over here,” and predicts the target encoder’s output for that region.

The loss is the distance between the predicted and actual target representations. Only the context encoder and predictor get gradients. The target encoder’s weights are an exponential moving average of the context encoder’s weights, which is the trick that keeps the representations from collapsing into something trivial.

Compare this to two other ways of learning from unlabeled images. Masked autoencoders hide patches and reconstruct their pixels, so the model spends a lot of effort on texture and fine detail that may not matter. Contrastive methods like SimCLR make two augmented views of the same image agree, which works well but depends heavily on hand-designed augmentations like crops and color jitter that encode assumptions about what should be ignored. I-JEPA reconstructs nothing and uses no augmentations. It predicts, in an abstract space, what’s in a part of the image it can’t see.

The results are good: strong linear-probe performance on ImageNet, better than masked autoencoders on some tasks that need semantic understanding, and much more efficient to train. They report training a huge ViT in under 72 hours on 16 A100s.

What I’m watching for is the step from images to video, and from video to action. Predicting a hidden patch of a still image teaches spatial structure. Predicting the representation of the next second of video, from the last few seconds, would teach something about how things move and interact, which is the actual world model part. And adding actions, predicting what the representation will be after I do something, is what would make it useful for planning. That’s the path from the paper, and I-JEPA is its first step.

My view from last year still holds. This is a good way to learn from observation, and I also think observation alone will only go so far. But learning to predict in representation space, and ignore unpredictable detail, is the right tool for building world models, whatever data they end up learning from.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…