Search papers, labs, and topics across Lattice.
LatentPress bypasses human-readable text and visual rendering by encoding conversational histories and long documents into continuous memory tokens injected directly into a frozen decoder's embedding layer without intermediate text reconstruction. This machine-facing compression paradigm circumvents decoding bottlenecks and KV-cache bloat, training only a tiny reader-matched adapter (4.2M–26.2M parameters, ~0.1% of the base model). Across benchmarks, it achieves 4–16× compression, reaching 0.504 accuracy on LongMemEval at 7.7× compression compared to 0.490 for raw uncompressed context, while speeding up context reading by 5–9× and writing by an order of magnitude over text summarization.
Language models process context more effectively through direct continuous embeddings than human-readable text, beating uncompressed context accuracy at 7.7× compression while slashing inference latency by up to 9×.
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference. A small reader-matched writer compresses $4$-$16\times$ while training only an adapter (4.2M-26.2M parameters, $\sim\!0.1\%$ of the decoder). On LongMemEval, LatentPress reaches $0.504$ accuracy at $7.70\times$ compression versus $0.490$ for uncompressed evidence, outperforming text summaries (0.184) and OCR-based compression (0.426 to 0.312). On LongBench-QA, in-domain writers match or exceed raw-context reading at $4$-$8\times$ compression, while $16\times$ trails raw. Writing takes 43ms per conversation, roughly an order of magnitude faster than text summarization or OCR reconstruction, and reading is $5$-$9\times$ faster than raw context or cached OCR. We validate the interface under two transfer settings, zero-shot from UltraChat to LongMemEval memory QA and from LongMemEval-derived QA to unseen LongBench document domains, establishing direct soft tokens as a practical machine-facing context interface beyond text and vision. The implementation of the experiments could be found at: https://github.com/HJSang/LatentPress .