Search papers, labs, and topics across Lattice.
This paper introduces IntroConformal, a training-free Conformal Risk Control framework that leverages introspective signals from large vision-language models (LVLMs) to ensure factual correctness in generated content. By employing layer-wise semantic stability and verification probability as conformity scores, the method provides finite-sample, distribution-free guarantees on factuality without relying on external verifiers. The results demonstrate that IntroConformal not only meets conformal risk guarantees but also significantly reduces abstention rates while outperforming existing baselines in claim-level discrimination.
Introspective signals from LVLMs can provide stronger factuality guarantees than traditional external verification methods, leading to fewer incorrect outputs.
Large Vision-Language Models (LVLMs) have achieved strong multimodal performance, yet ensuring the factual correctness of generated content remains challenging. Existing methods that provide statistical guarantees on factuality typically rely on external verifiers or generation-time confidence signals, which introduce auxiliary dependencies or often fail for confident but incorrect outputs. We argue that reliable factuality control can instead be achieved through introspective signals derived from the model itself. We introduce IntroConformal, a training-free Conformal Risk Control (CRC) framework that provides finite-sample, distribution-free factuality guarantees. We first instantiate it with layer-wise semantic stability, a conformity score derived from hidden-state representations, and then propose verification probability, a stronger score capturing the model's self-administered judgment on claim factuality. Across multiple LVLM architectures, IntroConformal satisfies the conformal risk guarantee while substantially reducing abstention and achieving competitive or superior claim-level discrimination relative to external verifier-based baselines.