Search papers, labs, and topics across Lattice.
4
0
8
2
ThoughtFold cuts token usage by 56% without sacrificing accuracy by folding reasoning chains and eliminating redundant explorations.
Depth Registers can slash W4A4 quantization error by nearly 14x, revealing critical insights into how model components contribute to performance degradation.
LLMs can edit code and text *much* faster by copying verbatim chunks from the input, achieving up to 303x speedup over autoregressive generation without end-to-end training.
Forget parallel probing – a commit-open protocol using SAE feature traces can reliably expose hosted LLM providers silently substituting cheaper models, even against adaptive attacks.