5 papers · 1 filter
FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
Yalun Wu, Junfeng Fang, Jiawei Wang +6
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close…
EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision
Chao Fei, Qingyi Si, Kaihua Liang +3
Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent sam…
A Self-Evolving Agent for Longitudinal Personal Health Management
Haoran Li, Jiebi Deng, Tong Jin +10
Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an open-source agent architecture…
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
Yue Pan, Ziheng Zhang, Junxiang Lei +3
Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This workflow creates a planning-laye…
When Agents Evolve, Institutions Follow
Chao Fei, Hongcheng Guo, Yanghua Xiao
Across millennia, complex societies have faced the same coordination problem of how to organize collective action among cognitively bounded and informationally incomplete individua…