20 citations · 46 across the 16 of their papers we have counts for
17 papers
MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
Bang Yang, Fenglin Liu, Xian Wu +3
Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for trainin…
Customizing General-Purpose Foundation Models for Medical Report Generation
Bang Yang, Asif Raza, Yuexian Zou +1
Medical caption prediction which can be regarded as a task of medical report generation (MRG), requires the automatic generation of coherent and accurate captions for the given med…
HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Dongchao Yang, Songxiang Liu, Rongjie Huang +3
Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations. Nowadays, audio codec models are increasingly…
TLAG: An Informative Trigger and Label-Aware Knowledge Guided Model for Dialogue-based Relation Extraction
Hao An, Dongsheng Chen, Weiyuan Xu +2
Dialogue-based Relation Extraction (DRE) aims to predict the relation type of argument pairs that are mentioned in dialogue. The latest trigger-enhanced methods propose trigger pre…
Improving Text-Audio Retrieval by Text-aware Attention Pooling and Prior Matrix Revised Loss
Yifei Xin, Dongchao Yang, Yuexian Zou
In text-audio retrieval (TAR) tasks, due to the heterogeneity of contents between text and audio, the semantic information contained in the text is only similar to certain frames w…
PoseRAC: Pose Saliency Transformer for Repetitive Action Counting
Ziyu Yao, Xuxin Cheng, Yuexian Zou
This paper presents a significant contribution to the field of repetitive action counting through the introduction of a new approach called Pose Saliency Representation. The propos…