Search papers, labs, and topics across Lattice.
Beihang University
1
0
3
Multimodal LLMs can cleanly disentangle camera ego-motion from real-world scene dynamics without any drone telemetry or pose sensors, using just four optical-flow-derived residual tokens per slice to unlock robust aerial video reasoning.