3 papers
cs.CV2026
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Wenzheng Zeng, Siyi Jiao, Chen Gao +2
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregres…
cs.CV2025
MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane Representation
Siyi Jiao, Wenzheng Zeng, Yerong Li +4
Human instance matting aims to estimate an alpha matte for each human instance in an image, which is challenging as it easily fails in complex cases requiring disentangling mingled…
cs.CV2024
DFIMat: Decoupled Flexible Interactive Matting in Multi-Person Scenarios
Siyi Jiao, Wenzheng Zeng, Changxin Gao +1
Interactive portrait matting refers to extracting the soft portrait from a given image that best meets the user's intent through their inputs. Existing methods often underperform i…