35 citations · 45 across the 12 of their papers we have counts for
12 papers
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
Jiliang Hu, Zuchao Li, Ping Wang +3
The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that i…
Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning
Cong Yang, Zuchao Li, Hongzan Jiao +2
Recently, while significant progress has been made in remote sensing image change captioning, existing methods fail to filter out areas unrelated to actual changes, making models s…
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
Xiangke Zeng, Zuchao Li, Lefei Zhang +3
Chinese Spelling Correction (CSC) stands as a foundational Natural Language Processing (NLP) task, which primarily focuses on the correction of erroneous characters in Chinese text…
BatGPT-Chem: A Foundation Large Model For Retrosynthesis Prediction
Yifei Yang, Runhan Shi, Zuchao Li +4
Retrosynthesis analysis is pivotal yet challenging in drug discovery and organic chemistry. Despite the proliferation of computational tools over the past decade, AI-based systems…
DSDRNet: Disentangling Representation and Reconstruct Network for Domain Generalization
Juncheng Yang, Zuchao Li, Shuai Xie +2
Domain generalization faces challenges due to the distribution shift between training and testing sets, and the presence of unseen target domains. Common solutions include domain a…
Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models
Juncheng Yang, Zuchao Li, Shuai Xie +3
Adapter-based parameter-efficient transfer learning has achieved exciting results in vision-language models. Traditional adapter methods often require training or fine-tuning, faci…