Search papers, labs, and topics across Lattice.
9
29
11
17
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
GroupVideo achieves unprecedented fidelity in multi-character video generation, resolving identity confusion and unnatural motions that plague existing methods.
WQ-Fusion achieves a remarkable score of 0.836 in cross-domain audio representation, showcasing the power of dynamic gated attention in feature selection.
A two-stage framework for mispronunciation detection in low-resource Arabic achieves a groundbreaking F1-score of 0.7201, outperforming previous methods by over 63%.
Adaptive WNG estimation via deep learning leads to significant gains in speech enhancement performance over conventional methods.
GraphPO slashes redundancy in reasoning model training, enabling more efficient exploration and improved performance on complex tasks.
SPRI achieves a remarkable 3.39 BLEU point improvement over the best existing MoE upcycling method, demonstrating that pretrained weight structures can be effectively leveraged for better expert diversity.
MemDreamer narrows the performance gap with human experts in long video understanding to just 3.7 points while processing only 2% of the full context.
GPT-4o now has open-source competition: Ming-Omni matches its modality support in a single, unified model capable of perception and generation across image, text, audio, and video.