3 papers
cs.AI2026
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
Hao Shen, Hang Yang, Zhouhong Gu +1
Large language models have advanced from single-turn question answering to deep research systems that iteratively decompose research questions, invoke retrieval tools, and synthesi…
cs.CR2026
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
Hao Shen, Zhouhong Gu, Haokai Hong +1
The widespread adoption of Large Language Models (LLMs) has raised significant privacy concerns regarding the exposure of personally identifiable information (PII) in user prompts.…
cs.AI2026
MARO: Learning Stronger Reasoning from Social Interaction
Yin Cai, Zhouhong Gu, Juntao Zhang +1
Humans face countless scenarios that require reasoning and judgment in daily life. However, existing large language model training methods primarily allow models to learn from exis…