Search papers, labs, and topics across Lattice.
This paper introduces LeVJEPA, a novel video encoder that utilizes a collapse-free objective to efficiently learn representations from video data without relying on complex architectural asymmetries or reconstruction methods. By employing a single encoder trained with an invariance loss and regularized by SIGReg, LeVJEPA achieves significant reductions in pretraining compute while improving downstream accuracy compared to existing methods. The results show that LeVJEPA can outperform traditional video pretraining approaches like V-JEPA 2 and DINOv2, making video a more viable substrate for general-purpose visual pretraining.
LeVJEPA achieves up to 20.8x less pretraining compute while surpassing the performance of leading video representation methods, reshaping the landscape of video-based learning.
Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coupling an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor, or circumvent it by reconstructing masked content in pixel space. We introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective, which dispenses with both. A single encoder is trained with an invariance loss over global and local views of a clip, regularized by SIGReg, which excludes collapse with a provable guarantee. The architecture reduces to an encoder and a projector, and the objective to a single hyperparameter. This formulation admits two properties. First, the cost of pretraining is governed by the number of tokens the encoder observes; uniform random token dropping renders this number small while simultaneously improving downstream accuracy. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute, and at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Second, since no asymmetry between branches is required, the encoder can be trained with block-causal attention at no measurable accuracy cost: temporal ordering becomes a property of the encoder itself. Against a compute-matched DINOv2 trained on frames of the same videos, LeVJEPA approaches the image-pretrained encoder on appearance-centric evaluation while nearly doubling its motion-centric accuracy. These results indicate that, once its computational overhead is removed, video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.