Search papers, labs, and topics across Lattice.
This paper introduces ExpertVerse, a benchmark designed to evaluate generative models on knowledge-intensive visual reasoning across nine cognitive capabilities and eight expert disciplines. By curating a dataset of 1,611 expert-annotated instances and developing the ExpertVerse-100K dataset, the authors facilitate a comprehensive assessment of reasoning in multimodal generative tasks. The training of KnowThinker, a reasoning engine utilizing a novel Bootstrapped Pareto Policy Optimization method, reveals significant reasoning deficits in existing models, underscoring the need for improved benchmarks in visual generation.
Existing generative models struggle with knowledge-intensive reasoning, but ExpertVerse reveals critical deficits that could redefine evaluation standards in multimodal AI.
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation. We develop \textbf{ExpertVerse}, a capability-centric benchmark to evaluate generative models via knowledge-intensive lens. ExpertVerse stratifies reasoning generation across an orthogonal taxonomy of \textit{9 cognitive capabilities} and \textit{8 expert disciplines}, yielding \textit{58 sub-disciplines}. We curate 1,611 expert-annotated instances covering single-image editing, multi-image composition, and text-to-image generation. We further develop an automated workflow to produce \textbf{ExpertVerse-100K}, a large-scale dataset with reasoning traces and knowledge-anchored rationale annotations. Based on this, we train \textbf{KnowThinker} with RL fine-tuning, a VLM reasoning engine with world knowledge that jointly generates thinking processes and refined instructions. Towards the cross-modal credit misalignment and multi-objective gradient conflicts in multi-reward optimization, we propose a tailored Bootstrapped Pareto Policy Optimization (BPPO), which synergizes Bootstrapping Reward Rectification (BRR) and Conflict-Aware Pareto Advantage Fusion (CPAF). Extensive results of both open-source and proprietary models exposes critical reasoning deficits, highlighting imperative for knowledge-intensive benchmarks towards next-generation visual generation.