Search papers, labs, and topics across Lattice.
This paper introduces a novel methodology for controlling refusal behavior in large language models (LLMs) through Stiefel-constrained rotation steering, leveraging Riemannian optimization for parameter-efficient transformations. The authors demonstrate that their approach outperforms existing methods reliant on auxiliary constructs, achieving greater intervention efficiency. An extensive ablation study underscores the significance of design choices in enhancing the effectiveness of the proposed steering mechanism.
Riemannian optimization enables a new level of control over LLM refusal behavior, outperforming traditional methods that depend on auxiliary constructs.
Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations. In our work, we develop a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization. We empirically validate the proposed scheme, demonstrating its superiority in intervention efficiency. An extensive ablation study highlights the importance of key design choices in our method. Our results identify the proposed rotation-based steering scheme as a promising direction for more reliable control over the behavior of LLMs.