9 papers · 1 filter
RAVE: Re-Allocating Visual Attention in Large Multimodal Models
Xi Leng, Xinhong Ma, Ziqiang Dong +4
Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allocation, including cross-moda…
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
Ran Ran, Jiwei Wei, Shuchang Zhou +5
Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query…
Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework
Ke Liu, Jiwei Wei, Shuchang Zhou +5
Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on generator-specific forgery p…
HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection
Shuchang Zhou, Kaiwen Shen, Jiwei Wei +3
The rapid evolution of generative models has enabled the creation of highly realistic and diverse synthetic images, posing significant challenges to reliable and generalizable Synt…
Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection
Shuchang Zhou, Shangkun Wu, Jiwei Wei +4
AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation Models (VFMs) provide rich s…
ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance
Yang Yang, Feifan Meng, Han Fang +1
Diffusion models have achieved remarkable success in image generation, yet their training is predominantly driven by full-reference objectives that enforce pixel-wise similarity to…