Shenzhen Technology UniversityShenzhen University of AdvancedMar 12, 2026arXiv:2603.11846

ZeroSense:How Vision matters in Long Context Compression

Yonghan Gao, Zehong Chen, Lijia Xu, Lijian Xu, Jingzhi Chen, Jingwei Guan, Xingyu Zeng

AI Summary

This paper introduces a new evaluation framework, ZeroSense, to assess the quality of visual-text compression (VTC) methods independently of downstream MLLM performance. ZeroSense uses a benchmark with low semantic correlation to ensure the evaluation reflects VTC quality, unaffected by MLLM inference. Experiments show a significant divergence between VTC quality measured by ZeroSense and downstream task accuracy, revealing limitations of existing evaluation protocols.

Key Contribution

Current methods for evaluating visual-text compression are misleading because they conflate compression quality with the downstream model's ability to fill in the gaps.

Abstract

Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs'capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating contextual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.

Computer Vision Eval Frameworks & Benchmarks Multimodal Models

Citation Metrics

Citations2

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

ZeroSense:How Vision matters in Long Context Compression

Related Papers