Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use
Zichen Tian, Jinpeng Chen, Cheng Gong +2
High-quality multi-turn tool-use data is essential for training agentic models, yet existing data synthesis methods often underrepresent the argument-level dependencies that are cr…
cs.AI2026
CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification
Jinpeng Chen, Cheng Gong, Hanbo Li +9
Developing multi-turn interactive tool-use agents is challenging because real-world user needs are often complex and ambiguous, yet agents must execute deterministic actions to sat…
cs.AI2025
Efficient Reasoning via Reward Model
Yuhao Wang, Xiaopeng Li, Cheng Gong +4
Reinforcement learning with verifiable rewards (RLVR) has been shown to enhance the reasoning capabilities of large language models (LLMs), enabling the development of large reason…