Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
Wei Wu, Liyi Chen, Congxi Xiao +7
Large reasoning models enhanced by reinforcement learning with verifiable rewards have achieved significant performance gains by extending their chain-of-thought. However, this par…
cs.AI2026
Knowledge-Graph Paths as Intermediate Supervision for Self-Evolving Search Agents
Huyu Wu, Jun Liu, Xiaochi Wei +3
Self-evolving search agents reduce reliance on human-written training questions by generating and solving their own search tasks. We build on Search Self-Play (SSP), a representati…
cs.AI2026
SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility
Xuyang Zhi, Peilun zhou, Chengqiang Lu +10
The evolution of Large Language Models (LLMs) is shifting the focus from single, verifiable tasks toward complex, open-ended real-world scenarios, imposing significant challenges o…