From the 1 of 14 linked papers with an AI index.
14 papers
MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
Zhuoning Xu, Xiucheng Zhang, Hanjun Luo +3
Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original authorization boundaries. Exist…
Evidence-Grounded AI for Musculoskeletal Care
Wenjie Li, Yujie Zhang, Fanrui Zhang +34
The paper presents OrthoPilot, a clinical AI system powered by a large language model that integrates real-time hospital data and external medical knowledge to provide evidence‑bas…
Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance
Peng Cui, Jitao Wang, Siyan Xue +40
Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging systems provide coarse categories, but often m…
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Hongcheng Gao, Hailong Qu, Jingyi Tang +18
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predomin…
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
Hanyu Li, Yichi Zhang, Speed Zhu +3
Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on end-to-end tasks such as issue…
Polar: Agentic RL on Any Harness at Scale
Binfeng Xu, Hao Zhang, Shaokun Zhang +9
Reinforcement learning for language agents increasingly depends on custom harnesses that manage long-running context, multi-turn tool use and multi-agent orchestration. However, po…