Search papers, labs, and topics across Lattice.
Thanks:
7
0
7
Bridging the gap between real and synthetic video representations can dramatically enhance zero-shot video captioning accuracy.
Transforming how LLM agents manage experiences, ToE achieves remarkable gains in problem-solving efficiency and accuracy, outpacing conventional methods.
TDVR achieves a remarkable 15.25% improvement in zero-shot 3D visual grounding accuracy by tackling ambiguous queries and viewpoint deficiencies head-on.
Achieve state-of-the-art micro-expression recognition with a Transformer-based architecture that slashes computational complexity without sacrificing accuracy.
Ditch slow, error-prone autoregressive video captioning: diffusion models can now generate captions in parallel, rivaling autoregressive quality with a significant speed boost.
Medical VQA gets a boost: Mamba-based models enhanced with knowledge graphs now outperform existing methods by better associating lesion features with diagnostic criteria.
LLMs can now capture an author's unique voice in translations, thanks to a multi-agent system guided by a "Stylistic Feature Spectrum" derived from wavelet transforms.