Search papers, labs, and topics across Lattice.
2
1
5
0
Token-level credit allocation using counterfactual replay boosts performance by an average of 4.4 percentage points without the need for an auxiliary scoring model.
Even the top-performing LLM struggles with cross-file reasoning, achieving only 69.1% accuracy on a new benchmark designed to reflect real-world software development challenges.