24 citations · 99 across the 67 of their papers we have counts for
63 papers · 1 filter
VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
Jiangang Zhu, Zheng Wang, Bin Zhu +2
Multi-expert models have become the dominant paradigm for long-tailed learning, largely attributed to their presumed ability to benefit from expert diversity. However, we revisit t…
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Yaole Wang, Xiaoyu Chen, Xin Ma +5
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…
SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection
Fei Li, Yue Yu, Yuran Wang +3
AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail to generalize to newly emergi…
DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection
Zihao Cai, Xinghan Li, Ruiyan Yang +3
As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge acquired from previous ones…
Disentangling Semantic Attention from Structural Bias in the Attention Manifold
Pengkun Jiao, Bin Zhu, Jingjing Chen +1
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…
Adaptive Inference-Time Scaling via Early-Step Latent Verification for Image Editing
Yue Yu, Yang Jiao, Jiayu Wang +2
Instruction-based image editing has made notable progress with recent advances in generative models. However, the quality of the edited result is still influenced by the randomly s…