From the 1 of 10 linked papers with an AI index.
1 citations · 1 across the 8 of their papers we have counts for
5 papers · 1 filter
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Senqiao Yang, Kaichen Zhang, Zhaoyang Jia +20
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process…
Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1
The paper introduces SportMV-Bench, a new benchmark for evaluating multimodal large language models on multi‑camera sports videos, and proposes SportMV-Agent, an agentic system tha…
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
Chuangchuang Tan, Xiang Ming, Jinglu Wang +5
The rapid advancement of AI-generated content (AIGC) has enabled the synthesis of visually convincing images; however, many such outputs exhibit subtle \textbf{semantic anomalies},…
ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
Chuangchuang Tan, Jinglu Wang, Xiang Ming +4
Advances in generative models have led to AI-generated images visually indistinguishable from authentic ones. Despite numerous studies on detecting AI-generated images with classif…