9 papers
ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement
Yufeng Yang, Jianzhuang Liu, Jisheng Chu +4
Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restricts their generalization ability a…
RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models
Yufeng Yang, Xianfang Zeng, Zhangqi Jiang +8
Image restoration under real-world degradations is critical for downstream tasks such as autonomous driving and object detection. However, existing restoration models are often lim…
MagicSeg: Open-World Segmentation Pretraining via Counterfactural Diffusion-Based Auto-Generation
Kaixin Cai, Pengzhen Ren, Jianhua Han +4
Open-world semantic segmentation presently relies significantly on extensive image-text pair datasets, which often suffer from a lack of fine-grained pixel annotations on sufficien…
TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
Yuqi Peng, Lingtao Zheng, Yufeng Yang +4
Personalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation…
DIVE: Taming DINO for Subject-Driven Video Editing
Yi Huang, Wei Xiong, He Zhang +4
Building on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and…
GLAD: Generalizable Tuning for Vision-Language Models
Yuqi Peng, Pengfei Wang, Jianzhuang Liu +1
Pre-trained vision-language models, such as CLIP, show impressive zero-shot recognition ability and can be easily transferred to specific downstream tasks via prompt tuning, even w…