3 papers
cs.SD2025
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
Jiliang Hu, Zuchao Li, Mengjia Shen +3
Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved gr…
cs.SD2024
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
Jiliang Hu, Zuchao Li, Ping Wang +3
The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that i…
cs.AI2024
Hypergraph based Understanding for Document Semantic Entity Recognition
Qiwei Li, Zuchao Li, Ping Wang +2
Semantic entity recognition is an important task in the field of visually-rich document understanding. It distinguishes the semantic types of text by analyzing the position relatio…