4 papers
Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
Jianing Geng, Ruiqi He, Zekun Fei +6
Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-…
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
Xinquan Chen, Zhenyun Yin, Shan He +38
As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…
Investigating CoT Monitorability in Large Reasoning Models
Shu Yang, Junchao Wu, Xilin Gong +4
Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex tasks by engaging in extended reasoning before producing final answers. Beyond improving abilities…
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
Shu Yang, Junchao Wu, Xuansheng Wu +3
Large Reasoning Models (LRMs) have achieved remarkable performance on complex tasks by engaging in extended reasoning before producing final answers, yet this strength introduces t…