Search papers, labs, and topics across Lattice.
5
0
10
0
Safety judgments for reasoning traces are far more complex than for final responses, revealing critical gaps in current guardrail models' capabilities.
Debugging near-miss hardware operators can yield a 66.7% success rate, outperforming traditional regeneration methods by a staggering margin.
SIRI allows LLM agents to autonomously develop and internalize skills, achieving up to a 2.2% performance boost without external dependencies.
A diffusion model can generate high-quality synthetic chromosome images, boosting anomaly detection by nearly 14% F1 score and reducing reliance on scarce real-world abnormal samples.
Mixing easy and hard examples during recommender pre-ranking hurts performance, but this new method disentangles them to boost user engagement by 0.4% in a real-world deployment.