Search papers, labs, and topics across Lattice.
7
1
9
3
Natural-language critiques can transform how we evaluate and optimize song generation models, leading to more human-aligned outputs.
Medical tabular data representation can be revolutionized by a framework that prioritizes feature importance and stabilizes ordinal relationships, setting a new state-of-the-art in multimodal pre-training.
Task-adaptive feature fusion can dramatically improve multi-task affective behavior analysis, achieving an overall validation score of 1.6341 on the ABAW11 benchmark.
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
SPACE achieves four to seventeen times fewer inter-robot collisions compared to traditional greedy planners, redefining efficiency in large-scale swarm exploration.
DINOv2 visual features and Wav2Vec 2.0 audio features can be effectively fused in a two-stage model to achieve state-of-the-art facial expression recognition in challenging, unconstrained video conditions.
Finally, a fully open-source, reproducible system for long-form song generation is here, complete with licensed data, code, and a Qwen-based model that rivals closed-source systems.