Search papers, labs, and topics across Lattice.
4
0
4
5
Internal representations can serve as powerful lie detectors, revealing discrepancies in LLM forecasts that CoT reasoning may obscure.
Transforming post-training from opaque reward optimization into a transparent process of auditing and sculpting the learning signal could revolutionize how we guide model behavior.
Steering neural networks through the intrinsic geometry of their activations unlocks more natural and controllable behaviors than traditional linear interventions.
LLMs often know the answer long before their "reasoning" suggests, wasting tokens on performative chain-of-thought.