2 papers
cs.AI2026
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
Can Wang, Haoran Chen, Haowen Gao +3
Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existin…
cs.CL2026
ZenGen: Social Mind for LLMs
ZenGen Team, Zing Team, Ao Xiang +57
As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track…