Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong, LLM Department, Tencent
3
0
7
0
Scaling native multimodal pre-training reveals that text-heavy data mixtures require larger models for optimal efficiency, challenging conventional resource allocation strategies.
Current AI agents only manage to complete 20.6% of complex real-world tasks, revealing a stark gap in their capabilities compared to human users.
A single model now rivals specialized vision-language models in understanding, while also generating and editing images, thanks to a unified discrete diffusion framework.