12 papers
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
Yunjia Qi, Zehua Yin, Xintong Shi +10
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to…
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists
Yuexi Yang, Alyssa Wu, Ji Luo +4
The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing ben…
Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?
Yangda Peng, Yunjia Qi, Hao Peng +11
Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliability of LaaJ for rubric sco…
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes
Jingru Chen, Yiming Liu, Mingtao Chen +5
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imp…
Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty
Jeonghyun Kim, SooKyung Kim, Richeng Xuan +1
The core of knowledge distillation lies in transferring the teacher's rich 'dark knowledge'-subtle probabilistic patterns that reveal how classes are related and the distribution o…
OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning
Seunghee Kim, Ingyu Bang, Seokgyu Jang +5
Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models…