Search papers, labs, and topics across Lattice.
The Chinese University of Hong Kong
2
0
4
VAD reveals that isolating visual evidence can dramatically improve target reconstruction in multimodal learning, leading to more accurate student outputs.
GUI agents struggle with semantically ambiguous actions, but HATS tackles this by iteratively exploring hard cases and refining instruction alignment, leading to significant performance gains.