Search papers, labs, and topics across Lattice.
6
0
7
6
Results show that the multi-agent design of RoboFind fits the demands of personalized object search, where verifying object identity before declaring completion is what makes the outcome something a user can rely on.
X$^2$Localizer boosts single-frame retrieval performance by over 4% while maintaining full-video accuracy, bridging the gap between evaluation benchmarks and real-world applications.
HOWTransfer achieves an 86% success rate in translating human hand motions into robot actions, surpassing teleoperation in user preference.
Fine-tuning vision-language models on a richly described dataset boosts their accuracy in interpreting subtle driver actions by 10 points, revealing a critical gap in existing training methods.
Faulty negatives and unreliable positives can be effectively managed in multi-modal learning, leading to significant improvements in driver distraction detection accuracy.
Correcting errors in long-video understanding doesn't have to be a nightmare: IMPACT-CYCLE slashes human arbitration costs by 4.8x while boosting VQA accuracy by intelligently decomposing the task and focusing human effort where it matters most.