Search papers, labs, and topics across Lattice.
14
0
16
24
Achieving high-quality audio reconstruction at unprecedented speeds, Qwen-Audio-VAE encodes 32 minutes of audio in just 541 ms.
Current AI agents only manage to complete 20.6% of complex real-world tasks, revealing a stark gap in their capabilities compared to human users.
EDA not only corrects the current memory write but also actively removes outdated information, leading to superior performance in long-context scenarios.
Verification of coding agent outputs is now the bottleneck, not generation, and targeted design can significantly enhance performance while curbing reward hacking.
Qwen-AgentWorld achieves unprecedented simulation fidelity, outperforming existing models and enabling scalable agentic reinforcement learning across diverse real-world environments.
Bypassing final-layer perturbations can significantly enhance reasoning capabilities in aligned LLMs, achieving better performance with zero memory overhead.
Qwen-RobotNav redefines navigation by allowing real-time reconfiguration of strategies, achieving unprecedented flexibility and performance across diverse tasks.
Qwen-RobotManip achieves a 20% relative improvement over the previous state-of-the-art in robotic manipulation, showcasing unprecedented generalization capabilities from diverse, open-source datasets.
Recursive composition of verifiable environments can boost reasoning performance in RL by up to 3.1 points while using only a fraction of the original environments.
MTP acceptance rates can be dramatically improved by addressing entropy fluctuations, leading to up to 1.8x faster RL training.
One model to control them all: Qwen-VLA achieves impressive zero-shot generalization across diverse robotic tasks and embodiments by unifying vision-language-action modeling.
Forget hand-crafted benchmarks: CUA-Gym's auto-generated training data lets computer-use agents crush existing open-source models on real-world tasks.
Despite achieving comparable overall scores, top-performing medical LLMs exhibit surprising differences in reasoning, evidence use, and longitudinal follow-up when evaluated on a new Chinese medical benchmark, revealing critical gaps in clinically actionable treatment planning.
LLM benchmark accuracy jumps 10% when evaluated on a cleaned-up version of Humanity's Last Exam, highlighting the significant impact of dataset noise on performance metrics.