collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

RAVE: Re-Allocating Visual Attention in Large Multimodal Models

Xi Leng, Xinhong Ma, Ziqiang Dong +4

Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allocation, including cross-moda…

cs.CV2026

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding

Ran Ran, Jiwei Wei, Shuchang Zhou +5

Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query…

cs.CV2026

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

Ke Liu, Jiwei Wei, Shuchang Zhou +5

Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on generator-specific forgery p…

cs.CV2026

HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection

Shuchang Zhou, Kaiwen Shen, Jiwei Wei +3

The rapid evolution of generative models has enabled the creation of highly realistic and diverse synthetic images, posing significant challenges to reliable and generalizable Synt…

cs.CV2026

Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection

Shuchang Zhou, Shangkun Wu, Jiwei Wei +4

AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation Models (VFMs) provide rich s…

cs.CV2026

ACPO: Anchor-Constrained Perceptual Optimization for Diffusion Models with No-Reference Quality Guidance

Yang Yang, Feifan Meng, Han Fang +1

Diffusion models have achieved remarkable success in image generation, yet their training is predominantly driven by full-reference objectives that enforce pixel-wise similarity to…