Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
StepOPSD: Step-Aware Online Preference Self-Distillation for Agent Reinforcement Learning
Yanfei Zhang, Xu Lin, Chenglin Wu
Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions…
cs.AI2025
Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning
Yanfei Zhang
Large Language Models (LLMs) have emerged as one of the most significant technological advancements in artificial intelligence in recent years. Their ability to understand, generat…