Search papers, labs, and topics across Lattice.
Thoughtworks
4
0
5
0
Sycophancy in language models isn't just a single flaw; it's a complex interplay of distinct modes that activate under different conditions.
Traditional benchmarks miss 82% of LLM performance, revealing a vast underestimation of their true capabilities in diverse tasks.
STRIDE reveals that training data influences can be efficiently traced in LLMs using sparse recovery, achieving attribution 13 times faster than traditional methods.
LLM activation spaces aren't linear, and exploiting their true geometry with "Curveball steering" unlocks more effective control than standard linear interventions.