659 citations · 835 across the 26 of their papers we have counts for
18 papers · 1 filter
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
Zhijian Liu, Zhuoyang Zhang, Samir Khaki +5
Semantic segmentation empowers numerous real-world applications, such as autonomous driving and augmented/mixed reality. These applications often operate on high-resolution images…
Fisher-aware Quantization for DETR Detectors with Critical-category Objectives
Huanrui Yang, Yafeng Huang, Zhen Dong +6
The impact of quantization on the overall performance of deep learning models is a well-studied problem. However, understanding and mitigating its effects on a more fine-grained le…
Magic-Me: Identity-Specific Video Customized Diffusion
Ze Ma, Daquan Zhou, Chun-Hsiao Yeh +6
Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven…
VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
Rongyu Zhang, Zefan Cai, Huanrui Yang +9
Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data point…
CVPR 2023 Text Guided Video Editing Competition
Jay Zhangjie Wu, Xiuyu Li, Difei Gao +17
Humans watch more than a billion hours of video per day. Most of this video was edited manually, which is a tedious process. However, AI-enabled video-generation and video-editing…
Large Language Models are Visual Reasoning Coordinators
Liangyu Chen, Bo Li, Sheng Shen +5
Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsen…