Search papers, labs, and topics across Lattice.
11
0
7
4
Achieving a 0.499 face similarity score, WithEveryone revolutionizes group image generation by ensuring identity preservation for up to ten individuals without direct face copying.
Rosetta achieves multimodal integration without catastrophic forgetting, enabling models to expand their capabilities while retaining prior knowledge.
GEAR accelerates image synthesis convergence by up to 10x while enhancing feature coherence, challenging the traditional decoupling of tokenizers and generators.
Compressing state-of-the-art image generation models by up to 75% without sacrificing quality could revolutionize resource efficiency in AI image synthesis.
CrossFlow achieves a remarkable 1.62 FID score with a single function evaluation, redefining the efficiency of image generation by eliminating the need for a separate decoder.
Frame-level causal attention is all you need for effective visual reconstruction in unified multimodal models.
Current image difference captioning benchmarks fail to capture semantic consistency and penalize hallucinations, but DiffCap-Bench offers a robust alternative that aligns with human expert judgments and predicts downstream utility for image editing.
LMMs can learn to generate images *and* improve their understanding abilities, without catastrophic forgetting, by carefully disentangling and sharing experts within a MoE architecture.
By tightly coupling reasoning, searching, and generation, Unify-Agent demonstrates that agent-based modeling can substantially improve world knowledge grounding in image synthesis, rivaling closed-source models.
Achieve SOTA in both visual generation and understanding by harmonizing generative and semantic representations within a single ViT architecture.
Ditch discrete visual tokens: UniCom achieves SOTA multimodal generation by compressing continuous semantic representations, unlocking better controllability and consistency in image editing.