Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
DR-Arena: an Automated Evaluation Framework for Deep Research Agents
Yiwen Gao, Ruochen Zhao, Yang Deng +1
As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task p…
cs.CL2025
Multilingual Agent-Based World Modeling for Social Science
Xuan Zhang, Wenxuan Zhang, Anxu Wang +2
Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interac…
cs.CL2025
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
Chi Seng Cheang, Hou Pong Chan, Wenxuan Zhang +1
Recent work suggests that LLMs "know what they don't know", positing that hallucinated and factually correct outputs arise from distinct internal processes and can therefore be dis…