2 citations · 2 across the 10 of their papers we have counts for
8 papers · 1 filter
Rethinking VLMs for Image Forgery Detection and Localization
Shaofeng Guo, Jiequan Cui, Richang Hong
With the rapid rise of Artificial Intelligence Generated Content (AIGC), image manipulation has become increasingly accessible, posing significant challenges for image forgery dete…
NOUS: Video-Driven 3D Human Reaction Generation via Observation-Reaction Mutual Steering
Yuan Zhou, Luanyuan Dai, Yongzhi Li +7
Video-driven 3D human reaction generation aims to synthesize 3D human motion in response to the action observed in a video, playing an important role in interactive multimedia syst…
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
Junbao Zhou, Yuan Zhou, Kesen Zhao +4
Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align…
DragNeXt: Rethinking Drag-Based Image Editing
Yuan Zhou, Junbao Zhou, Qingshan Xu +5
Drag-Based Image Editing (DBIE), which allows users to manipulate images by directly dragging objects within them, has recently attracted much attention from the community. However…
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
Yuan Zhou, Qingshan Xu, Jiequan Cui +4
Recently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed i…
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
Kaihang Pan, Zhaoyu Fan, Juncheng Li +6
The swift advancement in Multimodal LLMs (MLLMs) also presents significant challenges for effective knowledge editing. Current methods, including intrinsic knowledge editing and ex…