4 papers · 1 filter
TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces
Shu-Xun Yang, Cunxiang Wang, Haoke Zhang +12
Agentic systems augment large language models with external tools and iterative decision making, enabling complex tasks such as deep research, function calling, and coding. However…
DVD: A Robust Method for Detecting Variant Contamination in Large Language Model Evaluation
Renzhao Liang, Jingru Chen, Bo Jia +7
Evaluating large language models (LLMs) is increasingly confounded by \emph{variant contamination}: the training corpus contains semantically equivalent yet lexically or syntactica…
UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge
Yang Zhang, Cunxiang Wang, Lindong Wu +4
Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own.…
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
Shu-Xun Yang, Cunxiang Wang, Yidong Wang +3
Evaluating mathematical capabilities is critical for assessing the overall performance of large language models (LLMs). However, existing evaluation methods often focus solely on f…