activity
20142025
most citedCompressing Deep Convolutional Networks using Vector Quantization

1k citations · 1.1k across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

17 papers · 1 filter

cs.CV2023

Spatial Attention and Syntax Rule Enhanced Tree Decoder for Offine Handwritten Mathematical Expression Recognition

Zihao Lin, Jinrong Li, Fan Yang +4

Offline Handwritten Mathematical Expression Recognition (HMER) has been dramatically advanced recently by employing tree decoders as part of the encoder-decoder method. Despite the…

cs.CV20222 cited

Threshold-adaptive Unsupervised Focal Loss for Domain Adaptation of Semantic Segmentation

Weihao Yan, Yeqiang Qian, Chunxiang Wang +1

Semantic segmentation is an important task for intelligent vehicles to understand the environment. Current deep learning methods require large amounts of labeled data for training.…

cs.CV2022

Learning Visibility for Robust Dense Human Body Estimation

Chun-Han Yao, Jimei Yang, Duygu Ceylan +3

Estimating 3D human pose and shape from 2D images is a crucial yet challenging task. While prior methods with model-based representations can perform reasonably well on whole-body…

cs.CV20229 cited

3D Vision with Transformers: A Survey

Jean Lahoud, Jiale Cao, Fahad Shahbaz Khan +4

The success of the transformer architecture in natural language processing has recently triggered attention in the computer vision field. The transformer has been used as a replace…

cs.CV20221 cited

Automatically Discovering Novel Visual Categories with Self-supervised Prototype Learning

Lu Zhang, Lu Qi, Xu Yang +3

This paper tackles the problem of novel category discovery (NCD), which aims to discriminate unknown categories in large-scale image collections. The NCD task is challenging due to…

cs.CV20229 cited

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Rui Qian, Yeqing Li, Zheng Xu +3

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we ex…