3 papers
cs.CL2026
StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs
Youxin Zhu, Yixuan Ding, Peng Lai +3
Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, exi…
cs.CL2026
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
Yixia Li, Hongru Wang, Peng Lai +13
Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet m…
cs.AI2026
Anytime Safe PAC Efficient Reasoning
Chengyao Yu, Hao Zeng, Youxin Zhu +3
Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex tasks but suffer from high computational costs and latency. While selective thinking strategies im…