From the 1 of 5 linked papers with an AI index.
5 papers
Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1
The paper introduces SportMV-Bench, a new benchmark for evaluating multimodal large language models on multi‑camera sports videos, and proposes SportMV-Agent, an agentic system tha…
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
Kerui Chen, Jianrong Zhang, Ming Li +2
Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion…
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…
BVINet: Unlocking Blind Video Inpainting with Zero Annotations
Zhiliang Wu, Kerui Chen, Kun Li +2
Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusi…
Prompt-Aware Controllable Shadow Removal
Kerui Chen, Zhiliang Wu, Wenjin Hou +3
Shadow removal aims to restore the image content in shadowed regions. While deep learning-based methods have shown promising results, they still face key challenges: 1) uncontrolle…