Publications (16)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
Size Wu, Wenwei Zhang, Lumin Xu +4
Open-vocabulary dense prediction tasks including object detection and image segmentation have been advanced by the success of Contrastive Language-Image Pre-training (CLIP). CLIP m…
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
Jie Yang, Wang Zeng, Sheng Jin +5
The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle wit…
Whole-Body Human Pose Estimation in the Wild
Sheng Jin, Lumin Xu, Jin Xu +5
This paper investigates the task of 2D human whole-body pose estimation, which aims to localize dense landmarks on the entire human body including face, hands, body, and feet. As e…
UniFS: Universal Few-shot Instance Perception with Point Representations
Sheng Jin, Ruijie Yao, Lumin Xu +4
Instance perception tasks (object detection, instance segmentation, pose estimation, counting) play a key role in industrial applications of visual models. As supervised learning m…
GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
Ruijie Yao, Sheng Jin, Lumin Xu +5
Multi-Label Image Recognition (MLIR) is a challenging task that aims to predict multiple object labels in a single image while modeling the complex relationships between labels and…
CLIM: Contrastive Language-Image Mosaic for Region Representation
Size Wu, Wenwei Zhang, Lumin Xu +3
Detecting objects accurately from a large or open vocabulary necessitates the vision-language alignment on region representations. However, learning such a region-text alignment by…
ViPNAS: Efficient Video Pose Estimation via Neural Architecture Search
Lumin Xu, Yingda Guan, Sheng Jin +5
Human pose estimation has achieved significant progress in recent years. However, most of the recent methods focus on improving accuracy using complicated models and ignoring real-…
Ultra-High Resolution Segmentation via Boundary-Enhanced Patch-Merging Transformer
Haopeng Sun, Yingwei Zhang, Lumin Xu +2
Segmentation of ultra-high resolution (UHR) images is a critical task with numerous applications, yet it poses significant challenges due to high spatial resolution and rich fine d…
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
Cuifeng Shen, Lumin Xu, Xingguo Zhu +1
Video autoencoders compress videos into compact latent representations for efficient reconstruction, playing a vital role in enhancing the quality and efficiency of video generatio…
ZoomNAS: Searching for Whole-body Human Pose Estimation in the Wild
Lumin Xu, Sheng Jin, Wentao Liu +4
This paper investigates the task of 2D whole-body human pose estimation, which aims to localize dense landmarks on the entire human body including body, feet, face, and hands. We p…
F-LMM: Grounding Frozen Large Multimodal Models
Size Wu, Sheng Jin, Wenwei Zhang +4
Endowing Large Multimodal Models (LMMs) with visual grounding capability can significantly enhance AIs' understanding of the visual world and their interaction with humans. However…
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
Hao Zhang, Lumin Xu, Shenqi Lai +5
Current image-based keypoint detection methods for animal (including human) bodies and faces are generally divided into full-supervised and few-shot class-agnostic approaches. The…
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
Size Wu, Wenwei Zhang, Lumin Xu +6
Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations…
TCFormer: Visual Recognition via Token Clustering Transformer
Wang Zeng, Sheng Jin, Lumin Xu +5
Transformers are widely used in computer vision areas and have achieved remarkable success. Most state-of-the-art approaches split images into regular grids and represent each grid…
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
Jie Yang, Wang Zeng, Sheng Jin +4
Recent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pix…
Pose for Everything: Towards Category-Agnostic Pose Estimation
Lumin Xu, Sheng Jin, Wang Zeng +5
Existing works on 2D pose estimation mainly focus on a certain category, e.g. human, animal, and vehicle. However, there are lots of application scenarios that require detecting th…