activity
20182024
most citedCenterCLIP: Token Clustering for Efficient Text-Video Retrieval

128 citations · 240 across the 19 of their papers we have counts for

collaborators
Showing 2023Show all

14 papers · 1 filter

cs.CV20233 cited

DiverseMotion: Towards Diverse Human Motion Generation via Discrete Diffusion

Yunhong Lou, Linchao Zhu, Yaxiong Wang +2

We present DiverseMotion, a new approach for synthesizing high-quality human motions conditioned on textual descriptions while preserving motion diversity.Despite the recent signif…

cs.CV20233 cited

JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh Recovery

Jiahao Li, Zongxin Yang, Xiaohan Wang +3

In this study, we focus on the problem of 3D human mesh recovery from a single image under obscured conditions. Most state-of-the-art methods aim to improve 2D alignment technologi…

cs.CV20231 cited

Bird's-Eye-View Scene Graph for Vision-Language Navigation

Rui Liu, Xiaohan Wang, Wenguan Wang +1

Vision-language navigation (VLN), which entails an agent to navigate 3D environments following human instructions, has shown great advances. However, current agents are built upon…

cs.CV20233 cited

Clustering based Point Cloud Representation Learning for 3D Analysis

Tuo Feng, Wenguan Wang, Xiaohan Wang +2

Point cloud analysis (such as 3D segmentation and detection) is a challenging task, because of not only the irregular geometries of many millions of unordered points, but also the…

cs.CV2023

Kefa: A Knowledge Enhanced and Fine-grained Aligned Speaker for Navigation Instruction Generation

Haitian Zeng, Xiaohan Wang, Wenguan Wang +1

We introduce a novel speaker model \textsc{Kefa} for navigation instruction generation. The existing speaker models in Vision-and-Language Navigation suffer from the large domain g…

cs.CL2023

Whitening-based Contrastive Learning of Sentence Embeddings

Wenjie Zhuo, Yifan Sun, Xiaohan Wang +2

This paper presents a whitening-based contrastive learning method for sentence embedding learning (WhitenedCSE), which combines contrastive learning with a novel shuffled group whi…