80 citations · 180 across the 21 of their papers we have counts for
29 papers
CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image Segmentation
Zicheng Zhang, Yi Zhu, Jianzhuang Liu +2
Referring image segmentation aims at localizing all pixels of the visual objects described by a natural language sentence. Previous works learn to straightforwardly align the sente…
Learning Self-Regularized Adversarial Views for Self-Supervised Vision Transformers
Tao Tang, Changlin Li, Guangrun Wang +3
Automatic data augmentation (AutoAugment) strategies are indispensable in supervised data-efficient training protocols of vision transformers, and have led to state-of-the-art resu…
"My nose is running.""Are you also coughing?": Building A Medical Diagnosis Agent with Interpretable Inquiry Logics
Wenge Liu, Yi Cheng, Hao Wang +6
With the rise of telemedicine, the task of developing Dialogue Systems for Medical Diagnosis (DSMD) has received much attention in recent years. Different from early researches tha…
Automated Progressive Learning for Efficient Training of Vision Transformers
Changlin Li, Bohan Zhuang, Guangrun Wang +3
Recent advances in vision Transformers (ViTs) have come with a voracious appetite for computing power, high-lighting the urgent need to develop efficient training methods for ViTs.…
Image Comes Dancing with Collaborative Parsing-Flow Video Synthesis
Bowen Wu, Zhenyu Xie, Xiaodan Liang +3
Transferring human motion from a source to a target person poses great potential in computer vision and graphics applications. A crucial step is to manipulate sequential future mot…
DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
Changlin Li, Guangrun Wang, Bing Wang +3
Dynamic networks have shown their promising capability in reducing theoretical computation complexity by adapting their architectures to the input during inference. However, their…