Search papers, labs, and topics across Lattice.
4
0
5
Activation probes can predict future behaviors in reasoning models with up to 91% accuracy, enabling effective steering without sacrificing output quality.
Steering isn't just a trick; it's a fundamentally different way to adapt language models, offering localized, reversible control that traditional fine-tuning can't match.
Activation steering turns interpretability into a hands-on debugging tool, but watch out for unintended consequences and limited generalization.
Know where and by how much your PINN is wrong: a lightweight method provides pointwise error estimates without needing the true solution.