Search papers, labs, and topics across Lattice.
This study introduces Cross-Contextual Consistency (C3) as a novel metric for evaluating the credibility of large language models (LLMs) by assessing the stability of their responses across varied yet topic-aligned prompts. By analyzing 26 models across six benchmarks in reasoning, factuality, and code generation, the authors reveal that responses with minimal cross-contextual shifts correlate with higher accuracy and factual correctness. The findings suggest that C3 can serve as a valuable diagnostic tool for evaluating model performance, especially in saturated benchmark scenarios.
Answers that maintain stability across different contexts are significantly more likely to be correct, revealing a new lens for assessing LLM reliability.
Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic-aligned, content-neutral contextual variation. Building on this intuition, we operationalize Cross-Contextual Consistency (C3) by comparing model generations under original and perturbed prompts. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross-contextual shifts are more likely to be correct or factual. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered"saturate".