collaborators

7 papers

cs.CL2026

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

Yangfan Hu, Xuhan Tong, Haoyue Bai +5

Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or wh…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.AI2026

The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun +7

Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, interdependent action sequence…

cs.LG2026

Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models

Shuibai Zhang, Caspian Zhuang, Chihan Cui +8

Diffusion language models (DLMs) enable parallel, non-autoregressive text generation, yet existing DLM mixture-of-experts (MoE) models inherit token-choice (TC) routing from autore…

cs.SE2026

StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability

Haoyue Bai, Dong Wang, Long Chen +5

Large language model-based web agents have demonstrated strong performance on realistic web interaction tasks. However, existing evaluations are predominantly conducted under relat…

cs.LG2025

How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns

Haoyue Bai, Yiyou Sun, Wenjie Hu +5

Large Language Models (LLMs) display strikingly different generalization behaviors: supervised fine-tuning (SFT) often narrows capability, whereas reinforcement-learning (RL) tunin…