Search papers, labs, and topics across Lattice.
This paper explores the design of Mixture-of-Experts (MoE) architectures for vision encoders, revealing that fine-grained MoE topologies significantly enhance performance compared to both dense and standard MoE models. The authors introduce an auxiliary-loss-free balancing variant to optimize expert utilization and a specialized MoE kernel to reduce inference latency. Their largest MoE-ViE model achieves competitive zero-shot performance against a state-of-the-art encoder that is 1.7 times its size while operating at only 76% of the latency, outperforming all tested encoders on image and video benchmarks when aligned with a language model.
Fine-grained MoE designs can outperform dense vision encoders while dramatically reducing latency, challenging the status quo in image and video understanding.
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling and find that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts. We further propose an auxiliary-loss-free balancing variant for better expert utilization, and design a specialized MoE kernel to mitigate inference latency overhead. To enhance video capabilities while preserving image knowledge, we introduce frame-level distillation paired with a novel freezing mechanism. We pretrain a series of Mixture-of-Experts Vision Encoders (MoE-ViE) across a range of sizes, all consistently outperforming their dense counterparts. Our largest model matches the zero-shot performance of a SOTA encoder 1.7x its size at 76% of its latency. When aligned with an LLM, MoE-ViE surpasses all compared encoders on image and video benchmarks, including those with up to 5x more activated parameters. Code is available at https://github.com/facebookresearch/moe_vie.