collaborators

12 papers

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.CR2026

Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

Xu Li, Simon Yu, Minzhou Pan +5

LLM-based agents are becoming increasingly capable, yet their safety lags behind. This creates a gap between what agents can do and should do. This gap widens as agents engage in m…

cs.LG2026

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought

Jiachen Zhao, Yiyou Sun, Weiyan Shi +1

Large language models can generate long chain-of-thought (CoT) reasoning, yet prior work suggests that CoT can be post-hoc rationalization rather than a faithful reflection of the…

cs.AI2026

The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun +7

Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, interdependent action sequence…

cs.AI2026

Strategy Executability in Mathematical Reasoning: Leveraging Human-Model Differences for Effective Guidance

Weida Liang, Yiyou Sun, Shuyuan Nan +3

Example-based guidance is widely used to improve mathematical reasoning at inference time, yet its effectiveness is highly unstable across problems and models-even when the guidanc…

cs.AI2026

Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?

Yiyou Sun, Georgia Zhou, Haoyue Bai +4

Recent supervised fine-tuning (SFT) approaches have significantly improved language models' performance on mathematical reasoning tasks, even when models are trained at a small sca…