collaborators

12 papers

cs.AI2026

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

Yunjia Qi, Zehua Yin, Xintong Shi +10

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to…

cs.CL2026

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Jinyi Han, Yuanjian Xu, Ying Liao +6

Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evalu…

cs.SE2026

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists

Yuexi Yang, Alyssa Wu, Ji Luo +4

The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing ben…

cs.CL2026

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

Yiling Tao, Shihan Deng, Meiling Tao +3

Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval and reasoning to fulfill user…

cs.CL2026

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Yangda Peng, Yunjia Qi, Hao Peng +11

Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliability of LaaJ for rubric sco…

cs.CV2026

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

Jin Wang, Jianxiang Lu, Guangzheng Xu +10

Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video g…