large language models 2autonomous agents 1benchmarking 1contrastive learning 1long-horizon reasoning 1multi-step reasoning 1process evaluation 1reinforcement learning 1self-distillation 1
From the 2 of 14 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
Xingjian Wu, Xuhang Zhu, Xingchen Liu +6
The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…
cs.LG2026
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
Xingjian Wu, Junlin Liu, Xingchen Liu +6
The paper introduces Contrastive Reinforced Policy Optimization (CRPO), a method that frames on‑policy self‑distillation for large language models as a contrastive learning problem…