4 papers
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Lin Shi, Haowei Lin, Zixuan Zhu +123
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a…
ChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care Data
Jiangyuan Wang, Xuyong Chen, Junwei He +3
Long-horizon clinical simulation -- predicting how a patient's physiology evolves over years under specified interventions -- is central to chronic-disease care, yet existing elect…
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen +75
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
Jiaze Li, Haoran Xu, Shiding Zhu +2
The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains chal…