1 citations · 1 across the 14 of their papers we have counts for
10 papers · 1 filter
WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents
Hengrui Gu, Xiaotian Han, Kaixiong Zhou
Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute valid actions. A training traject…
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
Hengrui Gu, Xiaotian Han, Yujing Bian +2
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of large language models (LLMs), but it often suffers from \textit{restricted…
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
Shouren Wang, Wang Yang, Chuang Ma +7
Hybrid-thinking language models expose explicit /think and /no_think modes, but current designs do not separate them cleanly. Even in /no_think mode, models often emit long and sel…
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
Wang Yang, Debargha Ganguly, Xinpeng Li +5
Hybrid reasoning language models are commonly controlled through high-level Think/No-think instructions to regulate reasoning behavior, yet we found that such mode switching is lar…
All You Need is One: Capsule Prompt Tuning with a Single Vector
Yiyang Liu, James C. Liang, Heng Fan +7
Prompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning genera…
Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
Debargha Ganguly, Vikash Singh, Sreehari Sankar +7
Large language models (LLMs) show remarkable promise for democratizing automated reasoning by generating formal specifications. However, a fundamental tension exists: LLMs are prob…