Search papers, labs, and topics across Lattice.
This study addresses the challenge of automated novelty judgment in scientific research by identifying a systematic bias in large language models (LLMs) that leads to miscalibrated novelty assessments. The authors introduce the Think-Probe-Respond (TPR) method, which enhances LLM performance by probing latent novelty judgments during the reasoning phase and conditioning the final output on these insights. Results show that TPR improves novelty judgment accuracy by 22.30% and effectively reduces the tendency to classify ideas as "medium novel."
LLMs misjudge research ideas as "medium novel" due to a systematic bias, but a new probing method boosts their accuracy by over 22%.
Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as"medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent"medium novelty"bias.