Search papers, labs, and topics across Lattice.
Affiliation:
3
0
3
Across models from 124M to 7B parameters, token predictions collapse to a scale-invariant core of just 1% to 3% of the network, enabling direct, closed-form model editing in single units without gradient descent.
Transforming attention heads into interpretable feature detectors could revolutionize how we understand and trust transformer models.
A transformer with explicit fuzzy logic not only matches baseline performance but also reveals how it interprets grammatical structures, making model behavior legible.