22 citations · 64 across the 11 of their papers we have counts for
15 papers
Prototypical Transformer as Unified Motion Learners
Cheng Han, Yawen Lu, Guohao Sun +9
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoForme…
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
Guohao Sun, Can Qin, Huazhu Fu +2
Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medica…
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
Jiamian Wang, Guohao Sun, Pichao Wang +5
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, rel…
Parameter-Efficient Masking Networks
Yue Bai, Huan Wang, Xu Ma +3
A deeper network structure generally handles more complicated non-linearity and performs more competitively. Nowadays, advanced network designs often contain a large number of repe…
Dual Lottery Ticket Hypothesis
Yue Bai, Huan Wang, Zhiqiang Tao +2
Fully exploiting the learning capacity of neural networks requires overparameterized dense networks. On the other side, directly training sparse neural networks typically results i…
Adversarial Memory Networks for Action Prediction
Zhiqiang Tao, Yue Bai, Handong Zhao +3
Action prediction aims to infer the forthcoming human action with partially-observed videos, which is a challenging task due to the limited information underlying early observation…