Search papers, labs, and topics across Lattice.
3
0
7
3
Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks, establishing block-wise diffusion as a practical generative paradigm for multimodal GUI agents.
High-fidelity image synthesis does not require paired text from day one: pre-training visual priors on uncaptioned images before multimodal alignment beats conventional joint training pipelines to establish a new open-source DiT benchmark.
UI-Venus-2 achieves unprecedented environment coverage and task reliability, paving the way for dependable multimodal GUI agents in real-world applications.