activity
20192024
most citedMulti-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting

9 citations · 32 across the 11 of their papers we have counts for

collaborators

13 papers

cs.CV2024

WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization

Jiawei Ma, Yulei Niu, Shiyuan Huang +2

Language has been useful in extending the vision encoder to data from diverse distributions without empirical discovery in training domains. However, as the image description is mo…

cs.CV2023

Characterizing Video Question Answering with Sparsified Inputs

Shiyuan Huang, Robinson Piramuthu, Vicente Ordonez +2

In Video Question Answering, videos are often processed as a full-length sequence of frames to ensure minimal loss of information. Recent works have demonstrated evidence that spar…

cs.CV2023★ 3 cited

Supervised Masked Knowledge Distillation for Few-Shot Transformers

Han Lin, Guangxing Han, Jiawei Ma +3

Vision Transformers (ViTs) emerge to achieve impressive performance on many data-abundant computer vision tasks by capturing long-range dependencies among local features. However,…

cs.CV2023★ 2 cited

DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection

Jiawei Ma, Yulei Niu, Jincheng Xu +3

Generalized few-shot object detection aims to achieve precise detection on both base classes with abundant annotations and novel classes with limited training data. Existing approa…

cs.CV2022

TempCLR: Temporal Alignment Representation with Contrastive Learning

Yuncong Yang, Jiawei Ma, Shiyuan Huang +4

Video representation learning has been successful in video-text pre-training for zero-shot transfer, where each sentence is trained to be close to the paired video clips in a commo…

cs.CV2022

Video in 10 Bits: Few-Bit VideoQA for Efficiency and Privacy

Shiyuan Huang, Robinson Piramuthu, Shih-Fu Chang +1

In Video Question Answering (VideoQA), answering general questions about a video requires its visual information. Yet, video often contains redundant information irrelevant to the…