Search papers, labs, and topics across Lattice.
This study investigates the persistent underperformance of Chain-of-Thought models in pointwise document reranking compared to direct scoring models, confirming that the performance gap remains stable even with models up to 32 billion parameters. Through a series of stress tests involving reinforcement learning, fine-grained supervision, and architectural decoupling, the authors enhance classification accuracy and scores but find that the relative ranking gap remains unresolved. The results indicate that the inherent constraints of routing continuous relevance semantics through discrete text create a stable bottleneck in ranking signal resolution, challenging the notion that this issue can be easily remedied through targeted training interventions.
Despite enhancements in classification accuracy, the fundamental ranking gap in Chain-of-Thought models remains a stubborn bottleneck in pointwise reranking.
In pointwise document reranking, Chain-of-Thought models typically underperform direct scoring models. While existing diagnostics attribute this to inferior classification, score polarization, or calibration breakdown, whether targeted training can bridge this gap remains unclear. Our empirical study first confirms that this gap is stable across scales up to 32B parameters, ruling out model and data capacity confounders. We then apply stress tests utilizing reinforcement learning, fine-grained supervision, and architectural decoupling to explicitly repair these deviations. Although these interventions improve classification accuracy and absolute scores, the relative ranking gap persists. These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution, revealing a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.