Search papers, labs, and topics across Lattice.
This paper introduces Penelope, a novel framework that enhances structured reasoning in pretrained decoder-only Transformers by localizing recurrent computation within a selected decoder interval. By employing time-modulated GRU dynamics and a progressive curriculum that transfers visible reasoning into latent space, Penelope achieves competitive accuracy while significantly reducing inference latency compared to traditional methods. The findings demonstrate that localized latent refinement can effectively balance accuracy and efficiency, offering a promising alternative to increasing model size or chaining intermediate outputs.
Localized latent reasoning can slash inference latency while maintaining competitive accuracy, challenging the need for larger models or lengthy reasoning chains.
Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while reducing measured inference latency. These results show that latent refinement can be localized to a narrow decoder interval, reducing repeated full-decoder execution without generating a long visible reasoning trace and providing a practical accuracy-efficiency tradeoff for decoder-only Transformer models.