17 papers
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
Wang Yang, Debargha Ganguly, Xinpeng Li +5
Hybrid reasoning language models are commonly controlled through high-level Think/No-think instructions to regulate reasoning behavior, yet we found that such mode switching is lar…
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
Wang Yang, Hongye Jin, Shaochen Zhong +4
Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effortlessly process many originally exhaust…
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
Wang Yang, Zirui Liu, Hongye Jin +3
Recent language models exhibit strong reasoning capabilities, yet the influence of long-context capacity on reasoning remains underexplored. In this work, we hypothesize that curre…
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
Wang Yang, Xiang Yue, Vipin Chaudhary +1
Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still suffers from inefficient, overly lengthy…
WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents
Hengrui Gu, Xiaotian Han, Kaixiong Zhou
Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute valid actions. A training traject…
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
Hengrui Gu, Xiaotian Han, Yujing Bian +2
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of large language models (LLMs), but it often suffers from \textit{restricted…