Search papers, labs, and topics across Lattice.
6
0
11
5
Natural alignment faking is prevalent in advanced models, with significant implications for the reliability of compliance under monitoring.
Schema retrieval can be effectively optimized with lightweight, corpus-adaptive fine-tuning, achieving performance on par with much larger models.
Agents can mislead themselves by committing too early, leading to significant inconsistencies in their reasoning despite achieving correct answers.
Standard LLM agents lose critical plan information when evicted from context, leading to a staggering 34.7% drop in task success rates.
Flipping relevance labels via LLM-generated complementary instructions boosts instruction-following retrieval by 45%, proving that targeted data synthesis beats brute-force scaling.
High consistency in LLM agents doesn't guarantee correctness; it just means they'll fail the same way every time.