9 papers
ASI-Bench: At the Dawn of Artificial Superintelligence
Junwei Zhou, Zhen Sun, Binyu Li +39
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiab…
MemGym: a Long-Horizon Memory Environment for LLM Agents
Wujiang Xu, Yu Wang, Kai Mei +8
Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of personalized information in multi-…
Interactive Evaluation Requires a Design Science
Keyang Xuan, Peiyang Song, Pan Lu +10
AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other…
Auditing Agent Harness Safety
Chengzhi Liu, Yichen Guo, Yepeng Liu +8
LLM agents increasingly run inside execution harnesses that dispatch tools, allocate resources, and route messages between specialized components. However, a harness can return a c…
Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
Wenyue Hua, Tianyi Peng, Chi Wang +4
Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI systems evolve into autonomous agents…
Individual Turing Test: A Case Study of LLM-based Simulation Using Longitudinal Personal Data
Minghao Guo, Ziyi Ye, Wujiang Xu +3
Large Language Models (LLMs) have demonstrated remarkable human-like capabilities, yet their ability to replicate a specific individual remains under-explored. This paper presents…