Search papers, labs, and topics across Lattice.
KAIST
5
0
6
3
Compressing visual tokens in Omni-LLMs can cut input costs by over half while boosting accuracy beyond traditional methods.
SmellNet-V reveals that olfactory identities can be effectively paired with visual data, leading to a 7% improvement in smell classification accuracy.
Text dominance in Audio LLMs can be mitigated through a novel back-patching technique that enhances audio representations, challenging the status quo of multimodal processing.
Integrating trainable prompts into the audio encoder can significantly boost few-shot learning performance in Audio-Language Models, outperforming traditional text-only approaches.
Achieving robust audio-video generation without extensive training resources, this study reveals that a multi-verifier approach can dramatically enhance output quality across multiple dimensions.