activity
20242026
most citedEnd-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering

1 citations · 2 across the 3 of their papers we have counts for

collaborators

6 papers

cs.SD2026

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

Jiliang Hu, Wenfu Wang, Zuchao Li +6

Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks remain limited -- they are mainly Engli…

cs.SD20261 cited

End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering

Jiliang Hu, Zuchao Li, Baoyuan Qi +2

Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processi…

cs.SD20261 cited

Covo-Audio Technical Report

Wenfu Wang, Chenxing Li, Liqiang Zhang +23

In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…

cs.SD2026

SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away

Jiajia Li, Jiliang Hu, Ziyi Pan +4

Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce anci…

cs.SD2025

Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding

Jiliang Hu, Zuchao Li, Mengjia Shen +3

Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved gr…

cs.SD2024

VHASR: A Multimodal Speech Recognition System With Vision Hotwords

Jiliang Hu, Zuchao Li, Ping Wang +3

The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that i…