9 papers
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
Kerui Chen, Jianrong Zhang, Ming Li +2
Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion…
4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping
Xindan Zhang, Weilong Yan, Yufei Shi +5
Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing metho…
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO
Fufangchen Zhao, Songbai Tan, Xuerui Qiu +7
Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried in…
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…
DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection
Hongyuan Qi, Feifei Shao, Ming Li +2
The rapid evolution of video generation technologies poses a significant challenge to media forensics, as conventional detection methods often fail to generalize beyond their train…
Deepfake Detection Generalization with Diffusion Noise
Hongyuan Qi, Wenjin Hou, Hehe Fan +1
Deepfake detectors face growing challenges in generalization as new image synthesis techniques emerge. In particular, deepfakes generated by diffusion models are highly photorealis…