Search papers, labs, and topics across Lattice.
This paper introduces an interpretable adaptive sampling method for large language models (LLMs) that optimizes test-time scaling by dynamically adjusting the number of samples based on prompt complexity and model confidence. By employing a lightweight fuzzy controller, the approach allocates resources more effectively, assigning fewer samples to easier prompts and more to challenging ones, thus enhancing efficiency and interpretability in inference. Experimental results demonstrate that this adaptive method outperforms standard baselines while maintaining performance levels comparable to full-budget controls, indicating its potential for improving LLM reasoning efficiency.
Adaptive sampling can significantly enhance LLM reasoning efficiency by tailoring resource allocation based on prompt difficulty and model confidence.
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier or more confident prompts and more samples to harder or less certain prompts, making inference-time compute inspectable rather than fixed or opaque. We evaluate under a fair-alignment protocol with matched decoding settings and controlled answer selection, and compare against best-of-$N$, compute-aware scaling, and self-certainty-based baselines on question-answering and mathematical reasoning tasks. Across models and datasets, adaptive fuzzy control improves over several standard baselines and remains close to a selector-matched full-budget control while reducing the average number of samples. These findings suggest that interpretable adaptive sampling is a practical direction for more efficient test-time reasoning in large language models.