7 papers
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval
Yizhuo Xu, Chaojian Yu, Yuanjie Shao +3
Composed Image Retrieval (CIR) task aims to retrieve target images based on reference images and modification texts. Current CIR methods primarily rely on fine-tuning vision-langua…
Mutually Causal Semantic Distillation Network for Zero-Shot Learning
Shiming Chen, Shuhuang Chen, Guo-Sen Xie +1
Zero-shot learning (ZSL) aims to recognize the unseen classes in the open-world guided by the side-information (e.g., attributes). Its key task is how to infer the latent semantic…
VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models
Bowen Zheng, Yongli Xiang, Ziming Hong +4
Image-to-Video (I2V) generation models, which condition video generation on reference images, have shown emerging visual instruction-following capability, allowing certain visual c…
Prototype-Guided Curriculum Learning for Zero-Shot Learning
Lei Wang, Shiming Chen, Guo-Sen Xie +4
In Zero-Shot Learning (ZSL), embedding-based methods enable knowledge transfer from seen to unseen classes by learning a visual-semantic mapping from seen-class images to class-lev…
Few-Shot Object Detection via Spatial-Channel State Space Model
Zhimeng Xin, Tianxu Wu, Yixiong Zou +3
Due to the limited training samples in few-shot object detection (FSOD), we observe that current methods may struggle to accurately extract effective features from each channel. Sp…
Toward Realistic Camouflaged Object Detection: Benchmarks and Method
Zhimeng Xin, Tianxu Wu, Shiming Chen +5
Camouflaged object detection (COD) primarily relies on semantic or instance segmentation methods. While these methods have made significant advancements in identifying the contours…