works on

From the 2 of 20 linked papers with an AI index.

collaborators

20 papers

cs.LG2026

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Xingjian Wu, Xuhang Zhu, Xingchen Liu +6

The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…

cs.LG2026

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Xingjian Wu, Junlin Liu, Xingchen Liu +6

The paper introduces Contrastive Reinforced Policy Optimization (CRPO), a method that frames on‑policy self‑distillation for large language models as a contrastive learning problem…

cs.CL2026

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning

Wenxuan Jiang, Zining Fan, Zijian Zhang +6

Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creati…

cs.CL2026

ATLAS: All-round Testing of Long-context Abilities across Scales

Deli Huang, Cunguang Wang, Hongyin Tang +15

Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task family, masking two failure m…

cs.AI2026

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

Zhengkang Guo, Yiyang Li, Lin Qiu +7

As LLM-based agents increasingly rely on external tools, it is important to evaluate their ability to sustain tool-grounded reasoning beyond familiar workflows and short-range inte…

cs.AI2026

HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness

Jianing Wang, Linsen Guo, Zhengyu Chen +8

Recent advances in agentic harness with orchestration frameworks that coordinate multiple agents with memory, skills, and tool use have achieved remarkable success in complex reaso…