collaborators

5 papers

cs.AI2026

SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval

Ningyuan Li, Haiyang Shen, Mugeng Liu +4

Recent advances in large language models and tool-using agents have expanded the range of benchmarked web tasks. Yet an important class of specialized retrieval tasks remains under…

cs.AI2026

MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis

Haiyang Shen, Taian Guo, Xuanzhong Chen +11

Although LLMs have made substantial progress in reasoning, systematically producing frontier-level reasoning data remains difficult. Existing synthesis methods often have limited v…

cs.AI2026

DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

Sixiong Xie, Zhuofan Shi, Haiyang Shen +8

Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. F…

cs.CV2026

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

Zhuofan Shi, Peilun Jia, Baoqin Sun +4

Recent deep research systems have improved the ability of large language models to produce long, grounded reports through iterative retrieval and reasoning. However, most text-cent…

cs.AI2026

M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games

Sixiong Xie, Zhuofan Shi, Haiyang Shen +2

Existing benchmarks for LLM agents' social behavior typically focus on a single capability dimension and evaluate only behavioral outcomes, overlooking process signals from reasoni…