Search papers, labs, and topics across Lattice.
This paper introduces ConvergeFlow, a flow-based language model that ensures convergence to valid token embeddings by constraining the data predictor within the convex hull of token embeddings and training with a mean squared error objective. The method overcomes the limitations of existing continuous frameworks that rely on cross-entropy supervision, allowing for direct token prediction without a decoder. Experimental results on OpenWebText show that ConvergeFlow achieves competitive performance compared to both continuous and discrete language models, highlighting the efficacy of the flow-based approach in language modeling.
ConvergeFlow guarantees convergence to valid token embeddings, eliminating the need for cross-entropy supervision in flow-based language models.
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. However, existing continuous frameworks still rely on decoders supervised with cross entropy (CE) because the flow trajectories are not guaranteed to terminate at valid token embeddings. Motivated by this limitation, we introduce \textbf{ConvergeFlow}, an embedding-space flow-based LM, which constrains the data predictor to the convex hull of token embeddings and trains it solely with the mean squared error objective induced by flow matching. Under suitable regularity conditions, we prove that the resulting flow converges to valid token embeddings despite errors in the data predictor, enabling direct token prediction without a CE-supervised decoder. We further develop three sampling mechanisms for controlling the trade-off between the generative perplexity and entropy. Experiments on OpenWebText demonstrate that ConvergeFlow achieves performance competitive with existing continuous and discrete diffusion LMs. These findings demonstrate the potential of the flow-based paradigm for language modeling. Our code is available at https://github.com/Na-Li66/ConvergeFlow.