activity
20242026
most citedAdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning

1 citations · 2 across the 21 of their papers we have counts for

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

Grounded Scaling: Why Agentic AI Needs Deterministic Environments

Liang Ding, Xintong Wang

Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism , -step chain success degrades as . The AGI-t…

cs.AI2026

AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning

Yutong Wang, Siyuan Xiong, Xuebo Liu +4

While Multi-Agent Systems (MAS) excel in complex reasoning, they suffer from the cascading impact of erroneous information from individual agents. Current solutions often resort to…

cs.AI2026

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

Songlin Bai, Xintong Wang, Linlin Yu +12

In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every parameter must respect a regula…

cs.AI20261 cited

AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning

Liang Ding

Evaluating LLM agent trajectories is fundamentally task-specific: a code-debugging agent should be judged on Correctness and Error Handling, not on Fluency or Safety. Yet the domin…

cs.AI2026

AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling

Liang Ding

LLM-agent training pipelines routinely discard failed trajectories even though GPT-4o achieves only 14-20% on WebArena and below 55% pass@1 on ToolBench; even specialised systems a…

cs.AI2025

EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce

Rui Min, Zile Qiao, Ze Xu +18

Foundation agents have rapidly advanced in their ability to reason and interact with real environments, making the evaluation of their core capabilities increasingly important. Whi…