Search papers, labs, and topics across Lattice.
2
0
3
0
The findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification, and demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety.
Chain-of-thought is far more than surface imitation: LLMs carve out discrete, context-dependent geometric subspaces for functional operations like deduction and decomposition, with middle layers encoding the abstract reasoning step rather than the literal tokens.