Search papers, labs, and topics across Lattice.
Affiliation:
9
0
12
5
AVA-Encoder achieves a 73.1% relative improvement in video representation learning, enabling agents to produce cinematic-grade videos with far fewer resources.
Coding agents can achieve high accuracy on final specifications, yet 35% of them falter when faced with equivalent requirement histories, exposing a hidden vulnerability in their design.
Naive matching in diffusion distillation can inadvertently amplify errors due to hidden information in teacher models, leading to a surprising failure mode called Negative Branch Asymmetry.
Generators can dramatically improve their performance on long-tailed visual requests by leveraging a teach-then-search co-training approach, overcoming a critical knowledge boundary.
Current AI agents struggle with long-horizon professional tasks, achieving only 30% success in complex GUI workflows, revealing critical gaps in their capabilities.
Forget RL fine-tuning – RationalRewards unlocks latent image generation capabilities at test time simply by having the model critique and refine its own prompts.
LongCat-Next shatters the language-centric paradigm by unifying text, vision, and audio into a single autoregressive model with minimal modality-specific design, finally reconciling understanding and generation in discrete vision modeling.
A Qwen3-8B model, trained with a new SFT+RLAIF recipe on a challenging new benchmark, SWE-QA-Pro, beats GPT-4o in repository-level code understanding.
LLMs can be made better software engineers by pre-training them to reconstruct the messy, iterative development process that led to the final, clean code in repositories.