13 papers
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
Yunjia Qi, Zehua Yin, Xintong Shi +10
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to…
Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?
Yangda Peng, Yunjia Qi, Hao Peng +11
Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliability of LaaJ for rubric sco…
StoryAlign: Evaluating and Training Reward Models for Story Generation
Haotian Xia, Hao Peng, Yunjia Qi +4
Story generation aims to automatically produce coherent, structured, and engaging narratives. Although large language models (LLMs) have significantly advanced text generation, sto…
WildReward: Learning Reward Models from In-the-Wild Human Interactions
Hao Peng, Yunjia Qi, Xiaozhi Wang +3
Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deplo…
On the Paradoxical Interference between Instruction-Following and Task Solving
Yunjia Qi, Hao Peng, Xintong Shi +5
Instruction following aims to align Large Language Models (LLMs) with human intent by specifying explicit constraints on how tasks should be performed. However, we reveal a counter…
Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
Shiruo Hu, Wenbo Shan, Yingjia Li +16
Hydro-Science and Engineering (Hydro-SE) is a critical and irreplaceable domain that secures human water supply, generates clean hydropower energy, and mitigates flood and drought…