Search papers, labs, and topics across Lattice.
13
0
10
3
A single distilled model can outperform larger heterogeneous teachers by effectively integrating their strengths without interference.
Bridging disparate model families, Any-OPD achieves a 4.5% increase in performance while reducing model size by 80%.
JoyAI-Video-Edit achieves real-time video editing at 30 FPS while maintaining high fidelity and temporal consistency, outperforming traditional streaming editors.
Token-level explanations reveal how FEMRs leverage patient history, bridging the gap between black-box models and clinical trust.
AnchorEdit achieves state-of-the-art performance in multi-turn image editing by maintaining subject identity across 10+ interactions, revolutionizing iterative design workflows.
Ultra Flash achieves real-time high-resolution video generation at unprecedented frame rates, pushing the boundaries of what鈥檚 possible in streaming video AI.
Raw context outperforms compact memory designs, revealing that memory structure is crucial for effective video generation in action-conditioned models.
Multi-modal models fail to maintain coherent memory across video streams, revealing fundamental weaknesses in their memory architectures.
Real-time infinite video generation is now feasible, achieving over 1.3 million frames in a single 24-hour rollout without sacrificing quality.
Unlock the full potential of your pretrained video diffusion models with a surprisingly simple four-stage post-training framework that drastically improves visual quality, temporal coherence, and instruction following.
Linear transport flows between degraded and clean image domains enable fast, adaptable image restoration that outperforms existing methods in distortion-perception balance.
RL fine-tuning of hybrid autoregressive-diffusion models can be made significantly more stable and effective by averaging gradients across multiple diffusion trajectories and filtering autoregressive tokens for consistency.
Achieve real-time, synchronized audio-visual generation at 25 FPS by distilling a bidirectional diffusion model into a fast, autoregressive architecture, overcoming training instability with novel alignment and token handling techniques.