23 papers
SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection
Fei Li, Yue Yu, Yuran Wang +3
AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail to generalize to newly emergi…
Disentangling Semantic Attention from Structural Bias in the Attention Manifold
Pengkun Jiao, Bin Zhu, Jingjing Chen +1
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…
From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
Yue Yu, Jiayu Wang, Jiajia Shi +4
Makeup transfer aims to apply the makeup style of a reference portrait to a source portrait while preserving identity and background. Early methods formulate this task as unsupervi…
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
Ziyao Tang, Pengkun Jiao, Bin Zhu +3
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely…
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
Yian Li, Yang Jiao, Bin Zhu +4
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language…
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
Hui Zhang, Juntao Liu, Zongkai Liu +4
Instruction-based image editing aims to modify specific content within existing images according to user-provided instructions while preserving non-target regions. Beyond tradition…