Search papers, labs, and topics across Lattice.
9
0
15
4
Efficient context handling in video tasks can elevate multimodal models to new heights of agency and reasoning capability.
Even the best LLMs struggle with Olympiad-level combinatorics, achieving only 65.4% on a benchmark designed to expose their reasoning limitations.
Future-L1 shows that preserving visual semantics in latent space can dramatically enhance video event prediction accuracy, outperforming previous models by substantial margins.
Even state-of-the-art T2I models falter in generating scientifically accurate illustrations, revealing critical gaps in text-rendering and reasoning capabilities.
Smaller models can significantly enhance the training of larger models by providing structured exploration signals that improve performance without the noise of traditional methods.
By intelligently pruning attention heads based on their spatial or temporal roles and adaptively routing denoising steps through the network, PARE achieves significant computational savings in video generation without sacrificing quality.
Existing image-to-image evaluations miss a critical aspect: whether the output image actually preserves the content of the input.
Domain-aware federated learning: By creating separate prototype clusters for each domain, FedDAP significantly improves performance in heterogeneous federated learning scenarios where clients have different data distributions.
Pretraining isn't just about scaling data volume; daVinci-LLM's ablations reveal that data processing depth, domain-specific strategies, and compositional balance are equally critical for unlocking LLM capabilities.