Search papers, labs, and topics across Lattice.
Affiliation:
11
0
13
6
Manifold drift can lead to substantial misalignment in generative models, but ThermoDPO offers a powerful solution that anchors preference optimization to the pretrained data manifold.
A novel dense supervision strategy boosts image editing model performance by leveraging over 1,000 fine-grained concepts from a massive dataset of 12 million pairs.
A single poisoning phase can create a programmable backdoor in VLMs, enabling dynamic control over unseen target captions without retraining.
TBSM achieves state-of-the-art FID scores by transforming generative modeling through a novel scattering mechanism that enhances sample-level efficiency and reduces noise.
A novel framework achieves unprecedented dataset distillation speed and accuracy by directly minimizing information loss, setting a new benchmark in the field.
Language-action pretraining can lead to VLA policies that are not only more robust but also less dependent on visual cues, achieving up to 45% higher success rates in real-world tasks.
RODS synthesizes new training data on-the-fly, enabling agents to maintain high performance with 20x fewer trajectories than traditional methods.
Current MLLMs struggle with fine-grained spatial reasoning, achieving only 37.2 F1 on challenging tasks compared to human performance of 84.0 F1.
A single model now rivals specialized vision-language models in understanding, while also generating and editing images, thanks to a unified discrete diffusion framework.
Generative training not only enhances a model's ability to manipulate objects in images, but also surprisingly strengthens its spatial reasoning skills.
Forget GANs: APEX unlocks high-fidelity one-step image generation by extracting adversarial signals directly from the flow model itself, sidestepping the usual training instabilities.