Search papers, labs, and topics across Lattice.
Affiliation:
4
0
3
Relying on a single oracle for feedback can inflate perceived gains in LLM test generation by nearly 15 percentage points, masking the true effectiveness of evolution strategies.
Showing all visible security tests upfront boosts functional and security success rates by over 19% on average, but not all models benefit equally.
Code models may leverage tests more for semantic guidance than as executable specifications, with surprising implications for their performance consistency.
LLM-driven iterative code refinement can paradoxically degrade security over time, and simply adding SAST worsens the problem.