DAMOFudanShanghai InnovationShanghai Innovation InstitueJun 16, 2026arXiv:2606.18249

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Wujian Peng, Lingchen Meng, Yuxuan Cai, Xianwei Zhuang, Yuhuan Yang, Rongyao Fang, Chenfei Wu, Junyang Lin, Zuxuan Wu, Shuai Bai

AI Summary

This paper introduces UniAR, a unified autoregressive framework that utilizes a single discrete visual tokenizer to integrate visual understanding and generation, overcoming limitations of existing multimodal models that rely on separate tokenizers. By employing a pretrained vision encoder with multi-level feature fusion and a lookup-free bitwise quantization scheme, UniAR effectively scales the visual vocabulary while maintaining semantic and detail fidelity. The model achieves state-of-the-art results in image generation and editing, demonstrating significant advancements in multimodal understanding benchmarks through large-scale pre-training and reinforcement learning techniques.

Key Contribution

A single visual tokenizer in UniAR bridges the gap between understanding and generation, achieving state-of-the-art performance in image generation and editing.

Abstract

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. UniAR adapts a pretrained vision encoder with multi-level feature fusion and a lookup-free bitwise quantization scheme, preserving both high-level semantics and low-level details while scaling the effective visual vocabulary at minimal cost. Building on this, the unified autoregressive model adopts parallel-bitwise-prediction to jointly predict spatially grouped, multi-level visual codes, substantially reducing visual sequence length and accelerating generation. Finally, a diffusion-based visual decoder operates on discrete visual tokens to decode high-fidelity images. Through large-scale pre-training, followed by supervised fine-tuning and reinforcement learning, UniAR achieves state-of-the-art performance on image generation and image editing while remaining competitive on multimodal understanding benchmarks. The project page is available at https://sharelab-sii.github.io/uniar-web.

Multimodal Models

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Related Papers