activity
20242026
most citedHalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

4 citations · 5 across the 13 of their papers we have counts for

collaborators

17 papers

cs.CL2026

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

Zhaolu Kang, Yantao Liu, Tailong Luo +14

Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a stru…

cs.CL2026

Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

Zhaolu Kang, Meixin Wu, Yu Xue +6

Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such score…

cs.AI2026

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

Yunjia Qi, Zehua Yin, Xintong Shi +10

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to…

cs.SE2026

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists

Yuexi Yang, Alyssa Wu, Ji Luo +4

The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing ben…

cs.CL2026

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Yangda Peng, Yunjia Qi, Haotian Xia +8

Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliability of LaaJ for rubric sco…

cs.CV2026

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

Jingru Chen, Yiming Liu, Mingtao Chen +5

Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imp…