6 papers
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
Ran Ran, Jiwei Wei, Shuchang Zhou +5
Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query…
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
Ke Liu, Jiwei Wei, Shuchang Zhou +5
Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on generator-specific forgery p…
HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection
Shuchang Zhou, Kaiwen Shen, Jiwei Wei +3
The rapid evolution of generative models has enabled the creation of highly realistic and diverse synthetic images, posing significant challenges to reliable and generalizable Synt…
DiffLoRA: Generating Personalized Low-Rank Adaptation Weights with Diffusion
Yujia Wu, Yiming Shi, Jiwei Wei +3
Personalized text-to-image generation has gained significant attention for its capability to generate high-fidelity portraits of specific identities conditioned on user-defined pro…
Multimodality Invariant Learning for Multimedia-Based New Item Recommendation
Haoyue Bai, Le Wu, Min Hou +5
Multimedia-based recommendation provides personalized item suggestions by learning the content preferences of users. With the proliferation of digital devices and APPs, a huge numb…
MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete Representations
Heyuan Yao, Zhenhua Song, Yuyang Zhou +3
In this work, we present MoConVQ, a novel unified framework for physics-based motion control leveraging scalable discrete representations. Building upon vector quantized variationa…