4 citations · 4 across the 3 of their papers we have counts for
5 papers · 1 filter
Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1
Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid moti…
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
Kerui Chen, Jianrong Zhang, Ming Li +2
Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion…
BVINet: Unlocking Blind Video Inpainting with Zero Annotations
Zhiliang Wu, Kerui Chen, Kun Li +2
Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusi…
Prompt-Aware Controllable Shadow Removal
Kerui Chen, Zhiliang Wu, Wenjin Hou +3
Shadow removal aims to restore the image content in shadowed regions. While deep learning-based methods have shown promising results, they still face key challenges: 1) uncontrolle…