Search papers, labs, and topics across Lattice.
Affiliation:
4
0
4
This work proposes Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework that turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length.
Tailoring evaluation taxonomies for each vision-language model reveals a 32% improvement in performance and uncovers unique model blind spots that global assessments miss.
Achieving competitive video super-resolution quality with just 11.25% of the usual trainable parameters, LiteVSR redefines efficiency in adapting frozen diffusion models.
Smart glasses powered by web-native AI agents can now outperform commercial solutions in assistive tasks, offering a practical path to always-on, context-aware help for users navigating daily life.