9 papers
Disentangling 3D Modeling from Spatial Reasoning
Haoze Sun, Jiequan Cui, Qingshan Xu +1
In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perceptio…
Generalized Kullback-Leibler Divergence Loss
Jiequan Cui, Beier Zhu, Qingshan Xu +5
In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss…
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
Junbao Zhou, Yuan Zhou, Kesen Zhao +4
Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align…
MuSteerNet: Human Reaction Generation from Videos via Observation-Reaction Mutual Steering
Yuan Zhou, Yongzhi Li, Yanqi Dai +6
Video-driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for building human-like interactive AI…
Rethinking VLMs for Image Forgery Detection and Localization
Shaofeng Guo, Jiequan Cui, Richang Hong
With the rapid rise of Artificial Intelligence Generated Content (AIGC), image manipulation has become increasingly accessible, posing significant challenges for image forgery dete…
Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation
Kaichao Jiang, He Wang, Xiaoshuai Hao +5
Joint Energy-based Models (JEMs) are well known for their ability to unify classification and generation within a single framework. Despite their promising generative and discrimin…