Search papers, labs, and topics across Lattice.
Arena Intelligence Inc
9
0
12
Existing benchmarks miss the mark on faithfulness, but a new dependency-aware checklist reveals the true performance gaps in T2I models.
MIRAGE successfully turns the tables on image editing systems by leveraging their own moderation processes to block unauthorized manipulations.
DualEval reveals that unifying static and preference-based evaluations can lead to more reliable model rankings and deeper insights into item performance.
A VLM can autonomously evolve its questioning capabilities, producing harder and more diverse questions that enhance its overall performance without needing external data.
APEX reveals that optimizing data alongside prompts can boost LLM performance by over 11% while significantly reducing wasted compute resources.
Rethinking supervised fine-tuning as target distribution design reveals that optimizing token likelihood may overlook richer model knowledge, leading to significant performance gains.
One-Forcing achieves state-of-the-art one-step video generation while slashing training costs to a third of previous methods.
Forget training costly reward models for text-to-image alignment – AutoRubric-T2I learns interpretable rubrics that outperform them using less than 0.01% of the data.
Forget hand-crafted environments: ClawEnvKit lets you automatically generate diverse, verified environments for claw-like agents from natural language, slashing costs by 13,800x.