Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Kimi Team, Angang Du, Bofei Gao +93
Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learni…
cs.AI2025
SOP-Agent: Empower General Purpose AI Agent with Domain-Specific SOPs
Anbang Ye, Qianran Ma, Jia Chen +7
Despite significant advancements in general-purpose AI agents, several challenges still hinder their practical application in real-world scenarios. First, the limited planning capa…