7 papers
Attributing Structured-Output Gains in Function Calling: Interface Alignment versus Procedural Transfer
Wanyi Chen, Daoyuan Chen, Fang Kong
Structured-output benchmarks reward both task decisions and interface compliance, so prompt-induced function-calling gains require attribution before they can be interpreted as tra…
ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
Jie Zhu, Huaixia Dou, Shuo Jiang +5
Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited interpretability and little sup…
Harnesses for Inference-Time Alignment over Execution Trajectories
Boyuan Wang, Bochao Li, Minghan Wang +2
Harness engineering has emerged as an important inference-time technique for large language model (LLM) agents, aiming to improve long-term performance through task decomposition a…
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
Wanyi Chen, Xiao Yang, Xu Yang +7
We introduce Agent2 RL-Bench, a compact diagnostic benchmark for evaluating agentic RL post-training, which tests whether LLM agents can autonomously design, implement, debug, and…
SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems
Jialiang Fan, Weizhe Xu, Mengyu Liu +3
Safety-critical task planning in robotic systems remains challenging: classical planners suffer from poor scalability, Reinforcement Learning (RL)-based methods generalize poorly,…
CARE: Cognitive-reasoning Augmented Reinforcement for Emotional Support Conversation
Jie Zhu, Yuanchen Zhou, Shuo Jiang +5
Emotional Support Conversation (ESC) plays a vital role in alleviating psychological stress and providing emotional value through dialogue. While recent studies have largely focuse…