most citedExploiting Multi-modal Curriculum in Noisy Web Data for Large-scale Concept Learning

4 citations · 6 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV20241 cited

Efficient Training of Large Vision Models via Advanced Automated Progressive Learning

Changlin Li, Jiawei Zhang, Sihao Lin +4

The rapid advancements in Large Vision Models (LVMs), such as Vision Transformers (ViTs) and diffusion models, have led to an increasing demand for computational resources, resulti…

cs.CV2024

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

Xiaoyu Zhu, Hao Zhou, Pengfei Xing +6

In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel…

cs.CV20241 cited

Adversarially Masked Video Consistency for Unsupervised Domain Adaptation

Xiaoyu Zhu, Junwei Liang, Po-Yao Huang +1

We study the problem of unsupervised domain adaptation for egocentric videos. We propose a transformer-based model to learn class-discriminative and domain-invariant feature repres…

cs.CV20241 cited

ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Jiaming Zhou, Junwei Liang, Kun-Yu Lin +2

Zero-shot action recognition (ZSAR) aims to learn an alignment model between videos and class descriptions of seen actions that is transferable to unseen actions. The text queries…

cs.CV2023

Spatial-Temporal Alignment Network for Action Recognition

Jinhui Ye, Junwei Liang

This paper studies introducing viewpoint invariant feature representations in existing action recognition architecture. Despite significant progress in action recognition, efficien…

cs.CV20222 cited

Multi-dataset Training of Transformers for Robust Action Recognition

Junwei Liang, Enwei Zhang, Jun Zhang +1

We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Al…