19 citations · 21 across the 6 of their papers we have counts for
5 papers · 1 filter
C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Rongchang Li, Zhenhua Feng, Tianyang Xu +5
Compositional actions consist of dynamic (verbs) and static (objects) concepts. Humans can easily recognize unseen compositions using the learned concepts. For machines, solving su…
Investigating Self-Supervised Methods for Label-Efficient Learning
Srinivasa Rao Nandam, Sara Atito, Zhenhua Feng +2
Vision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks like classification…
LT-ViT: A Vision Transformer for multi-label Chest X-ray classification
Umar Marikkar, Sara Atito, Muhammad Awais +1
Vision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). Howev…
SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-supervised Skeleton-based Action Recognition
Cong Wu, Xiao-Jun Wu, Josef Kittler +4
Contrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal re…
Masked Momentum Contrastive Learning for Zero-shot Semantic Understanding
Jiantao Wu, Shentong Mo, Muhammad Awais +3
Self-supervised pretraining (SSP) has emerged as a popular technique in machine learning, enabling the extraction of meaningful feature representations without labelled data. In th…