Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning

Publication Date: 7/8/2026

Event: The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

Reference: pp. 27250–27268

Authors: Renliang Sun, University of California, Los Angeles; Wei Cheng, NEC Laboratories America, Inc.; Dawei Li, Arizona State University; Wei Wang, University of California, Los Angeles; Haifeng Chen, NEC Laboratories America, Inc.

Abstract: Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning — so-called overthinking— can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN (REFlective-Redundancy for Adaptive INference), a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking. REFRAIN integrates a two-stage stop discriminator to identify reflective yet redundant reasoning and a sliding-window Upper Confidence Bound (SW-UCB) multi-armed bandit controller to dynamically adjust stopping thresholds according to problem difficulty without supervision or fine-tuning. Across four representative benchmarks and two model families, REFRAIN re-duces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting. Extensive ablation and robustness analyses demonstrate its stability across models, scorers, and prompt variations. In summary, our findings highlight when-to-stop as a new and practical axis of test-time scaling — enabling models to reason not just more, but just enough.

Publication Link: https://aclanthology.org/2026.acl-long.1256/

0 replies

Leave a Reply

Want to join the discussion?
Feel free to contribute!

Leave a Reply