20 papers
GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization
Haozhen Yan, Siyuan Shan, Zijian Yu +4
AI-generated image manipulation localization identifies edited pixels, but its OOD performance lags behind image-level detection partly because pixel supervision entangles forensic…
Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision
Beibei Zhang, Chao Xu, Jun Lan +4
Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise i…
Maintain Plasticity in Long-timescale Continual Test-time Adaptation
Yanshuo Wang, Xuesong Li, Jinguang Tong +5
Continual test-time domain adaptation (CTTA) aims to adjust pre-trained source models to perform well over time across non-stationary target environments. While previous methods ha…
RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
Yan Hong, Wei Li, Kedong Xiu +6
Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online d…
LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models
Yan Hong, Kedong Xiu, Wei Li +6
We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving gener…
COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
Haozhen Yan, Yan Hong, Jiahui Zhan +5
Recent advances in image manipulation have enabled highly photorealistic content generation, but also lowered the barrier to arbitrary editing, raising concerns about multimedia au…