large language models 3autonomous agents 1benchmarking 1contrastive learning 1feedback-driven policy discovery 1generative recommendation 1intent modeling 1knowledge distillation 1long-horizon reasoning 1multi-step reasoning 1online inference 1process evaluation 1
From the 3 of 11 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
Xingjian Wu, Xuhang Zhu, Xingchen Liu +6
The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…
cs.LG2026
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
Xingjian Wu, Junlin Liu, Xingchen Liu +6
The paper introduces Contrastive Reinforced Policy Optimization (CRPO), a method that frames on‑policy self‑distillation for large language models as a contrastive learning problem…