1 citations · 1 across the 3 of their papers we have counts for
5 papers · 1 filter
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
Jiliang Hu, Zuchao Li, Ping Wang +3
The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that i…
Multi-modal Auto-regressive Modeling via Visual Words
Tianshuo Peng, Zuchao Li, Lefei Zhang +3
Large Language Models (LLMs), benefiting from the auto-regressive modelling approach performed on massive unannotated texts corpora, demonstrates powerful perceptual and reasoning…
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
Xiangke Zeng, Zuchao Li, Lefei Zhang +3
Chinese Spelling Correction (CSC) stands as a foundational Natural Language Processing (NLP) task, which primarily focuses on the correction of erroneous characters in Chinese text…
Hypergraph based Understanding for Document Semantic Entity Recognition
Qiwei Li, Zuchao Li, Ping Wang +2
Semantic entity recognition is an important task in the field of visually-rich document understanding. It distinguishes the semantic types of text by analyzing the position relatio…
The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
Jiajia Li, Lu Yang, Mingni Tang +4
Benchmark plays a pivotal role in assessing the advancements of large language models (LLMs). While numerous benchmarks have been proposed to evaluate LLMs' capabilities, there is…