Search papers, labs, and topics across Lattice.
4
0
7
54
Most developers edit AI-generated code within 15 minutes, often discarding the original completions entirely, highlighting a critical gap in LLM training data.
LLM agents can identify reproducibility problems in 90% of analyzed machine learning papers, leveraging GitHub issues as a novel supervision source.
Even GPT-5 only achieves 63% accuracy on time series anomaly questions from real software incidents, but a model-expert combination reaches 87%, highlighting the potential for hybrid intelligence in incident response.
Multimodal agents still struggle with game development, solving only ~50% of tasks in a new benchmark, GameDevBench, highlighting the need for better multimodal reasoning in complex software environments.