works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.CV2026

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Senqiao Yang, Kaichen Zhang, Zhaoyang Jia +20

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process…

cs.AI2026

Coordinated Networking for On-Device Agent-Augmented Real-Time Communication

Goodsol Lee, Juheon Yi, Jinglu Wang +3

AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents autonomously retrieve, analyze,…

cs.CV2026

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1

The paper introduces SportMV-Bench, a new benchmark for evaluating multimodal large language models on multi‑camera sports videos, and proposes SportMV-Agent, an agentic system tha…

cs.AI2026

Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft

Juheon Yi, Jinglu Wang, Xiaoyi Zhang +1

We present TickingCollabBench, a Minecraft-based multi-agent benchmark for a novel class of time-sensitive complementary collaboration tasks. Our benchmark reflects four core chara…

cs.CV2026

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

Kerui Chen, Jinglu Wang, Jianrong Zhang +3

Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…

cs.CL2026

A Decade-Scale Benchmark Evaluating LLMs' Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations

Andong Tan, Shuyu Dai, Jinglu Wang +7

Clinical practice guidelines (CPGs) play a pivotal role in ensuring evidence-based decision-making and improving patient outcomes. While Large Language Models (LLMs) are increasing…