collaborators

5 papers

cs.CL2026

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

Yiwen Gao, Ruochen Zhao, Yang Deng +1

As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task p…

cs.AI2026

Self-Evolving World Models for LLM Agent Planning

Xuan Zhang, Wenxuan Zhang, See-Kiong Ng +1

World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignor…

cs.AI2026

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond

Meng Chu, Xuan Billy Zhang, Kevin Qinghong Lin +47

As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics becomes a central bottleneck. Agents that ma…

cs.CL2026

Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness

Chi Seng Cheang, Hou Pong Chan, Wenxuan Zhang +1

Recent work suggests that LLMs "know what they don't know", positing that hallucinated and factually correct outputs arise from distinct internal processes and can therefore be dis…

cs.CL2025

Multilingual Agent-Based World Modeling for Social Science

Xuan Zhang, Wenxuan Zhang, Anxu Wang +2

Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interac…