Search papers, labs, and topics across Lattice.
Affiliation:
3
0
3
WorldVIEW, a multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings, is introduced, and the revision layer itself is identified as a previously undocumented, causal source of this stereotyping.
Switching an LLM from an API to a consumer chat interface can degrade performance more than downgrading an entire model generation鈥攁nd standard API hyperparameter controls cannot reliably bridge the gap.
Over a third of real-world information-seeking queries are high-risk, exposing critical failures in LLM responses, especially for analytical tasks.