37 citations · 40 across the 6 of their papers we have counts for
6 papers
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
Runze He, Kai Ma, Linjiang Huang +6
Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose F…
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
Hongyu Li, Tianrui Hui, Zihan Ding +5
Panoptic narrative grounding (PNG), whose core target is fine-grained image-text alignment, requires a panoptic segmentation of referred objects given a narrative caption. Previous…
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
Lijia Lv, Weigang Zhang, Xuehai Tang +4
Jailbreak vulnerabilities in Large Language Models (LLMs) refer to methods that extract malicious content from the model by carefully crafting prompts or suffixes, which has garner…
Learning to Discover Forgery Cues for Face Forgery Detection
Jiahe Tian, Peng Chen, Cai Yu +4
Locating manipulation maps, i.e., pixel-level annotation of forgery cues, is crucial for providing interpretable detection results in face forgery detection. Related learning objec…
InfoCSE: Information-aggregated Contrastive Learning of Sentence Embeddings
Xing Wu, Chaochen Gao, Zijia Lin +3
Contrastive learning has been extensively studied in sentence embedding learning, which assumes that the embeddings of different views of the same sentence are closer. The constrai…
RaP: Redundancy-aware Video-language Pre-training for Text-Video Retrieval
Xing Wu, Chaochen Gao, Zijia Lin +3
Video language pre-training methods have mainly adopted sparse sampling techniques to alleviate the temporal redundancy of videos. Though effective, sparse sampling still suffers i…