9 papers
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework…
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
Wei-Chieh Huang, Henry Peng Zou, Yaozu Wu +12
Deep research frameworks have shown promising capabilities in synthesizing comprehensive reports from web sources. While deep research possesses significant potential to address co…
Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
Yankai Chen, Xinni Zhang, Yifei Zhang +6
Brain-Computer Interfaces (BCIs) offer a direct communication pathway between the human brain and external devices, holding significant promise for individuals with severe neurolog…
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
Yaozu Wu, Jizhou Guo, Dongyuan Li +9
Effective guardrails are essential for safely deploying LLM-based agents in critical applications. Despite recent advances, existing guardrails suffer from two fundamental limitati…
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu +10
Recent improvements in large language models (LLMs) have led many researchers to focus on building fully autonomous AI agents. This position paper questions whether this approach i…
TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency
Henry Peng Zou, Zhengyao Gu, Yue Zhou +7
Test-time computing approaches, which leverage additional computational resources during inference, have been proven effective in enhancing large language model performance. This w…