Search papers, labs, and topics across Lattice.
2
0
6
0
Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks, establishing block-wise diffusion as a practical generative paradigm for multimodal GUI agents.
High-fidelity image synthesis does not require paired text from day one: pre-training visual priors on uncaptioned images before multimodal alignment beats conventional joint training pipelines to establish a new open-source DiT benchmark.