Search papers, labs, and topics across Lattice.
Affiliation:
6
0
7
7
Output format can distort perceived model capabilities, with a 40-point accuracy gain in one format vanishing in another.
Conditioned training outperforms traditional mixed joint training, revealing that the structure of prompts can significantly impact multimodal document understanding performance.
Relying on a single oracle for feedback can inflate perceived gains in LLM test generation by nearly 15 percentage points, masking the true effectiveness of evolution strategies.
Showing all visible security tests upfront boosts functional and security success rates by over 19% on average, but not all models benefit equally.
Code models may leverage tests more for semantic guidance than as executable specifications, with surprising implications for their performance consistency.
GRPO fails to improve performance in small language models, revealing a critical scale-dependent limitation in reinforcement learning applications.