Search papers, labs, and topics across Lattice.
5
0
8
54
Closing the capability gap is essential for realizing a future where drones revolutionize logistics and emergency response on a national scale.
Captions selected with VEGAS align significantly better with human attention, boosting retrieval performance and challenging the status quo of video captioning metrics.
Robots can now anticipate and adapt to environmental changes, revealing route failures that traditional planning methods miss.
Forget adversarial training: a closed-form solution can make multi-agent RL for drone collision avoidance surprisingly robust to GPS spoofing.
Ditch BLEU and ROUGE: ViSIL offers a unified metric for multimodal video captioning that actually correlates with VQA performance and human judgment by measuring information loss via VLM inference.