23 papers
HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
Xiao Zhang, Yusheng Wang, Yuhao Fei +5
Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delaye…
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design
Zejun Liu, Jian Wu, Ru Peng +4
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focus…
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework…
PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
Weiwei Ye, Hangchen Liu, Dongyuan Li +1
Large language models have become capable reasoners and tool users that write and run code and search the literature, which makes automating the research process itself a realistic…
TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning
Hangchen Liu, Dongyuan Li, Renhe Jiang +3
Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires both numerical accuracy and…
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu +17
Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-based agents still face significant…