Search papers, labs, and topics across Lattice.
4
0
8
4
Across widely adopted grounding benchmarks and navigation benchmarks spanning multiple platforms, LLaDA-UI substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks, establishing block-wise diffusion as a practical generative paradigm for multimodal GUI agents.
High-fidelity image synthesis does not require paired text from day one: pre-training visual priors on uncaptioned images before multimodal alignment beats conventional joint training pipelines to establish a new open-source DiT benchmark.
A novel dense supervision strategy boosts image editing model performance by leveraging over 1,000 fine-grained concepts from a massive dataset of 12 million pairs.
A single model now rivals specialized vision-language models in understanding, while also generating and editing images, thanks to a unified discrete diffusion framework.