Search papers, labs, and topics across Lattice.
Fraunhofer Heinrich Hertz Institute, Technological University Dublin
7
0
9
The results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.
A relevance-based concept atlas reveals how tissue morphology directly influences spatial transcriptomics predictions, enhancing interpretability in pathology.
Activation probes can predict future behaviors in reasoning models with up to 91% accuracy, enabling effective steering without sacrificing output quality.
Steering isn't just a trick; it's a fundamentally different way to adapt language models, offering localized, reversible control that traditional fine-tuning can't match.
Activation steering turns interpretability into a hands-on debugging tool, but watch out for unintended consequences and limited generalization.
Know where and by how much your PINN is wrong: a lightweight method provides pointwise error estimates without needing the true solution.