works on

From the 4 of 18 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Xingjian Wu, Xuhang Zhu, Xingchen Liu +6

The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…

cs.LG2026

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Xingjian Wu, Junlin Liu, Xingchen Liu +6

The paper introduces Contrastive Reinforced Policy Optimization (CRPO), a method that frames on‑policy self‑distillation for large language models as a contrastive learning problem…

cs.LG2026

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

Tao Chen, Gangwei Jiang, Pengyu Cheng +10

Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) pipelines. However, current rew…

cs.LG2026

Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works

Wenhua Nie, Jianan Wu, Junlin Liu +6

Group Relative Policy Optimization (GRPO) is a standard algorithm for reinforcement learning from verifiable rewards, but its group-mean-centered advantage can fail under binary re…

cs.LG2026

The Coupling Tax: How Shared Token Budgets Undermine Visible Chain-of-Thought Under Fixed Output Limits

Wenhua Nie, Junlin Liu, Jianan Wu +5

Chain-of-thought reasoning is often treated as a monotone way to improve language-model accuracy by letting a model think longer. We identify a countervailing effect, the coupling…

cs.LG2026

SciTS: Scientific Time Series Understanding and Generation with LLMs

Wen Wu, Ziyang Zhang, Liwei Liu +12

The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents u…