8 papers
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Weihuang Zheng, Tianyuan Zou, Eileen Ye +5
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, a…
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
Shanhui Zhao, Jiacheng Liu, Guohong Liu +13
AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span applications and data s…
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
Litian Liu, Reza Pourreza, Yubing Jian +2
Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection method…
Unfolding an Atomistic World: Atomistic Simulation of Reactor Pressure Vessel Steel Across Year-and-Meter Scales
Haozhi Han, Ruge Zhang, Haoquan Chen +8
Lifetime prediction of reactor pressure vessel (RPV) steel requires bridging atomistic degradation mechanisms with service-scale spatial and temporal regimes, from Angstroms and pi…
Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies
Binghan Wu, Shoufeng Wang, Yunxin Liu +3
The evolution toward Level 4 (L4) Autonomous Networks (AN) represents a strategic inflection point in telecommunications, where networks must transcend reactive automation to achie…
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
Hao Wen, Yifan Su, Feifei Zhang +4
Recent advances in Large Language Models (LLMs) have been driven by test-time compute scaling - a strategy that improves reasoning by generating longer, sequential thought processe…