Search papers, labs, and topics across Lattice.
13
0
11
13
MedUAG sets a new standard in medical multimodal models, achieving strong performance across diverse understanding and generation tasks with the largest dataset yet.
DentAgent outperforms senior specialists by 17.3 percentage points in multi-label diagnosis, revolutionizing multimodal dental reasoning with traceable evidence integration.
MedUP reveals that integrating visual perception and language understanding in a single token space can dramatically enhance performance in medical vision-language tasks.
The first publicly available dataset for early pregnancy fetal ultrasound screening could revolutionize automated diagnostics and standardization in prenatal care.
ClinFusion sets a new benchmark in multimodal medical understanding, outperforming both open-source and proprietary models in generating clinically relevant insights from diverse medical images.
Synthetic data that looks good can still tank your model's performance – Optimsyn uses influence functions to find the *actually* useful synthetic examples and optimize your generation rubrics.
Get clinically-accurate 3D dental models from a single panoramic X-ray, slashing radiation exposure and cost.
Vision-language models falter at the fine-grained temporal recognition crucial for surgical video understanding, while SurgRec excels.
LLM agents can achieve 3x faster web search and higher accuracy by dynamically routing between multiple context management strategies.
Forget real-world video datasets: training VLMs on just 7.7K synthetic videos with temporal primitives beats 165K real-world examples, unlocking surprisingly effective transfer learning for video reasoning.
Directly modeling 3D geometry in dental scans unlocks a 9.58% accuracy boost in multi-disease diagnosis compared to methods relying on 2D or multi-view image representations.
Group chats can be revitalized with LLM-powered agents, boosting message volume by nearly 30% in real-world deployments.
Current video benchmarks are too simple; UniVBench offers the first unified framework to measure the integrated capabilities of video foundation models using complex, multi-shot videos and a standardized evaluation system.