Search papers, labs, and topics across Lattice.
6
0
8
67
Unlocking the potential of unused fine-tuning data can enhance reasoning models' performance in complex domains without sacrificing their core capabilities.
Generative compilation enables AI models to receive compiler feedback during code generation, drastically reducing errors and improving code quality in real-time.
Modern embedding models excel in general IR tasks but falter in complex mathematical domains, revealing a critical gap in current evaluation benchmarks.
LLMs can follow detailed code refactoring instructions, but still fall short of mimicking human refactoring choices in real-world codebases, highlighting a critical gap in their ability to autonomously improve code quality.
LLM benchmark translations can be dramatically improved by test-time compute scaling, revealing a surprisingly cheap way to get more reliable multilingual evaluations.
Context files like AGENTS.md, intended to guide coding agents, often *hurt* performance and increase costs, challenging the common practice of using them.