16 papers
Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label
Cheng Chen, Yifan Zhao, Jia Li
Learning multi-label image classification with incomplete annotations is a challenging task that has been widely studied for its superior trade-off between high efficiency and less…
To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations
Wenzhuang Wang, Yifan Zhao, Mingcan Ma +4
Camouflage image generation (CIG) focuses on generating visually concealed objects that seamlessly blend into their backgrounds. Existing methods typically follow either background…
Adapting Vision-Language Models from Iconic to Inclusive for Multi-Label Recognition Without Labels
Cheng Chen, Jingyu Zhou, Yifan Zhao +1
Understanding multi-label images remains a challenging task in computer vision. With the rapid progress of vision-language multimodal learning, vision-language models (VLMs) enable…
Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation
Nan Bao, Yifan Zhao, Wenzhuang Wang +1
The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and dist…
Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Events
Yunshan Qi, Lin Zhu, Nan Bao +2
Novel view synthesis from low dynamic range (LDR) blurry images, which are common in the wild, struggles to recover high dynamic range (HDR) and sharp 3D representations in extreme…
Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL
Ruitao Wu, Yifan Zhao, Guangyao Chen +1
Few-Shot Class-Incremental Learning (FSCIL) challenges models to sequentially learn new classes from minimal examples without forgetting prior knowledge, a task complicated by the…