11 papers
Coordinated Networking for On-Device Agent-Augmented Real-Time Communication
Goodsol Lee, Juheon Yi, Jinglu Wang +3
AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents autonomously retrieve, analyze,…
Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1
The paper introduces SportMV-Bench, a new benchmark for evaluating multimodal large language models on multi‑camera sports videos, and proposes SportMV-Agent, an agentic system tha…
Closed-Loop Triplet Synergistic Generation for Long-Form Video
Xinlei Yin, Xiulian Peng, Xiao Li +2
Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllabil…
Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft
Juheon Yi, Jinglu Wang, Xiaoyi Zhang +1
We present TickingCollabBench, a Minecraft-based multi-agent benchmark for a novel class of time-sensitive complementary collaboration tasks. Our benchmark reflects four core chara…
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
Yuexi Du, Jinglu Wang, Shujie Liu +2
Large visual language models (VLMs) have shown strong multi-modal medical reasoning ability, but most operate as end-to-end black boxes, diverging from clinicians' evidence-based,…
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
Peng Xia, Jinglu Wang, Yibo Peng +10
Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across div…