collaborators

8 papers

cs.AI2026

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

Weihuang Zheng, Tianyuan Zou, Eileen Ye +5

Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, a…

cs.AI2026

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction

Shanhui Zhao, Jiacheng Liu, Guohong Liu +13

AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span applications and data s…

cs.AI2026

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

Litian Liu, Reza Pourreza, Yubing Jian +2

Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection method…

cs.DC2026

Unfolding an Atomistic World: Atomistic Simulation of Reactor Pressure Vessel Steel Across Year-and-Meter Scales

Haozhi Han, Ruge Zhang, Haoquan Chen +8

Lifetime prediction of reactor pressure vessel (RPV) steel requires bridging atomistic degradation mechanisms with service-scale spatial and temporal regimes, from Angstroms and pi…

cs.AI2026

Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies

Binghan Wu, Shoufeng Wang, Yunxin Liu +3

The evolution toward Level 4 (L4) Autonomous Networks (AN) represents a strategic inflection point in telecommunications, where networks must transcend reactive automation to achie…

cs.CL2025

ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute

Hao Wen, Yifan Su, Feifei Zhang +4

Recent advances in Large Language Models (LLMs) have been driven by test-time compute scaling - a strategy that improves reasoning by generating longer, sequential thought processe…