10 citations · 13 across the 4 of their papers we have counts for
4 papers
Associating Spatially-Consistent Grouping with Text-supervised Semantic Segmentation
Yabo Zhang, Zihao Wang, Jun Hao Liew +4
In this work, we investigate performing semantic segmentation solely through the training on image-sentence pairs. Due to the lack of dense annotations, existing text-supervised me…
Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring
Ruyang Liu, Jingjia Huang, Ge Li +3
Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention f…
Temporal Perceiving Video-Language Pre-training
Fan Ma, Xiaojie Jin, Heng Wang +4
Video-Language Pre-training models have recently significantly improved various multi-modal downstream tasks. Previous dominant works mainly adopt contrastive learning to achieve g…
Knowledge Guided Bidirectional Attention Network for Human-Object Interaction Detection
Jingjia Huang, Baixiang Yang
Human Object Interaction (HOI) detection is a challenging task that requires to distinguish the interaction between a human-object pair. Attention based relation parsing is a popul…