4 papers · 1 filter
The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory
Yixiong Chen, Xinyi Bai, Alan Yuille
Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environme…
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories
Yixiong Chen, Alan Yuille
Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitatio…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Meissa: Multi-modal Medical Agentic Intelligence
Yixiong Chen, Xinyi Bai, Yue Pan +2
Multi-modal large language models (MM-LLMs) have shown strong performance in medical image understanding and clinical reasoning. Recent medical agent systems extend them with tool…