Search papers, labs, and topics across Lattice.
8
0
11
13
MeanFlowNFT achieves superior few-step generation performance, outperforming traditional multi-step RL-tuned methods while maintaining efficiency.
Predictive divergence masks can significantly enhance the stability of RL updates in LLMs, outperforming traditional methods by aligning direction criteria with actual divergence changes.
Multimodal unlearning could revolutionize how we handle sensitive data in AI, enabling targeted removal without sacrificing model performance.
Flow-DPPO outperforms traditional PPO methods by achieving higher rewards and greater training stability through a novel divergence proximal constraint.
Smooth gradient adjustments in DRPO prevent harmful policy shifts, leading to more stable and efficient LLM training.
Forget slow, complex training: you can now distill diffusion models to just 4 steps and still beat the state-of-the-art in preference alignment, aesthetics, and composition.
OpenSearch-VL offers a fully transparent recipe for training state-of-the-art multimodal search agents, finally democratizing access to a capability previously locked behind closed doors.
Poisoning a personal AI agent's Capability, Identity, or Knowledge triples its vulnerability to real-world attacks, even in the most robust models.