activity
20232025
most citedSALMONN: Towards Generic Hearing Abilities for Large Language Models

20 citations · 32 across the 14 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2025

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

Mengqi Wang, Zhan Liu, Zengrui Jin +3

Diffusion-based large language models (DLLMs) have recently attracted growing interest as an alternative to autoregressive decoders. In this work, we present an empirical study on…

eess.AS2024

SOT Triggered Neural Clustering for Speaker Attributed ASR

Xianrui Zheng, Guangzhi Sun, Chao Zhang +1

This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be ap…

eess.AS2023★ 5 cited

Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models

Guangzhi Sun, Wenyi Yu, Changli Tang +6

Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but…

eess.AS2023★ 2 cited

Conditional Diffusion Model for Target Speaker Extraction

Theodor Nguyen, Guangzhi Sun, Xianrui Zheng +2

We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuou…

eess.AS2023

Connecting Speech Encoder and Large Language Model for ASR

Wenyi Yu, Changli Tang, Guangzhi Sun +6

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies a…