activity
20142024
most citedDenseNet: Implementing Efficient ConvNet Descriptor Pyramids

659 citations · 835 across the 26 of their papers we have counts for

collaborators
Showing cs.CVShow all

18 papers · 1 filter

cs.CV2024

Sparse Refinement for Efficient High-Resolution Semantic Segmentation

Zhijian Liu, Zhuoyang Zhang, Samir Khaki +5

Semantic segmentation empowers numerous real-world applications, such as autonomous driving and augmented/mixed reality. These applications often operate on high-resolution images…

cs.CV2024

Fisher-aware Quantization for DETR Detectors with Critical-category Objectives

Huanrui Yang, Yafeng Huang, Zhen Dong +6

The impact of quantization on the overall performance of deep learning models is a well-studied problem. However, understanding and mitigating its effects on a more fine-grained le…

cs.CV20243 cited

Magic-Me: Identity-Specific Video Customized Diffusion

Ze Ma, Daquan Zhou, Chun-Hsiao Yeh +6

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven…

cs.CV2024

VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness

Rongyu Zhang, Zefan Cai, Huanrui Yang +9

Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data point…

cs.CV20236 cited

CVPR 2023 Text Guided Video Editing Competition

Jay Zhangjie Wu, Xiuyu Li, Difei Gao +17

Humans watch more than a billion hours of video per day. Most of this video was edited manually, which is a tedious process. However, AI-enabled video-generation and video-editing…

cs.CV202314 cited

Large Language Models are Visual Reasoning Coordinators

Liangyu Chen, Bo Li, Sheng Shen +5

Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsen…