collaborators

9 papers

cs.CV2024

FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction

Runze He, Kai Ma, Linjiang Huang +6

Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose F…

cs.CV2024

Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding

Hongyu Li, Tianrui Hui, Zihan Ding +5

Panoptic narrative grounding (PNG), whose core target is fine-grained image-text alignment, requires a panoptic segmentation of referred objects given a narrative caption. Previous…

cs.CV2024

Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection

Cai Yu, Shan Jia, Xiaomeng Fu +6

With the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific…

cs.CV2024

Model Will Tell: Training Membership Inference for Diffusion Models

Xiaomeng Fu, Xi Wang, Qiao Li +3

Diffusion models pose risks of privacy breaches and copyright disputes, primarily stemming from the potential utilization of unauthorized data during the training phase. The Traini…

cs.CV2023

OSM-Net: One-to-Many One-shot Talking Head Generation with Spontaneous Head Motions

Jin Liu, Xi Wang, Xiaomeng Fu +4

One-shot talking head generation has no explicit head movement reference, thus it is difficult to generate talking heads with head motions. Some existing works only edit the mouth…

cs.CV2023

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Shaofei Huang, Han Li, Yuqing Wang +5

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semant…