Search papers, labs, and topics across Lattice.
William & Mary
3
0
3
MLLMs exhibit a staggering 80.22% bias towards incorrect outputs when faced with repeated UI patterns, revealing a critical flaw in their code generation capabilities.
Naive statistical analyses can lead to false positives in software engineering experiments, as demonstrated by the overestimation of prompt engineering's impact on code generation when confounding bias isn't addressed.
Forget what you know about prompt engineering: more specific system prompts don't always improve code generation, and few-shot examples can actually *hurt* performance in large code models.