From the 1 of 11 linked papers with an AI index.
7 papers · 1 filter
Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1
The paper introduces SportMV-Bench, a new benchmark for evaluating multimodal large language models on multi‑camera sports videos, and proposes SportMV-Agent, an agentic system tha…
Closed-Loop Triplet Synergistic Generation for Long-Form Video
Xinlei Yin, Xiulian Peng, Xiao Li +2
Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllabil…
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
Chuangchuang Tan, Xiang Ming, Jinglu Wang +5
The rapid advancement of AI-generated content (AIGC) has enabled the synthesis of visually convincing images; however, many such outputs exhibit subtle \textbf{semantic anomalies},…
ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
Chuangchuang Tan, Jinglu Wang, Xiang Ming +4
Advances in generative models have led to AI-generated images visually indistinguishable from authentic ones. Despite numerous studies on detecting AI-generated images with classif…
GS-Marker: Generalizable and Robust Watermarking for 3D Gaussian Splatting
Lijiang Li, Jinglu Wang, Xiang Ming +1
In the Generative AI era, safeguarding 3D models has become increasingly urgent. While invisible watermarking is well-established for 2D images with encoder-decoder frameworks, gen…
StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams
Yang LI, Jinglu Wang, Lei Chu +4
The advent of 3D Gaussian Splatting (3DGS) has advanced 3D scene reconstruction and novel view synthesis. With the growing interest of interactive applications that need immediate…