most citedLLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

2 citations · 4 across the 7 of their papers we have counts for

collaborators

11 papers

cs.CL2025

IG-Pruning: Input-Guided Block Pruning for Large Language Models

Kangyu Qiao, Shaolei Zhang, Yang Feng

With the growing computational demands of large language models (LLMs), efficient inference has become increasingly critical for practical deployment. Depth pruning has emerged as…

cs.CL2025

AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment

Mengyu Bu, Shaolei Zhang, Zhongjun He +2

Multilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities. However, their performance and cross-lingual alignment often la…

cs.LG2025

PSO-Merging: Merging Models Based on Particle Swarm Optimization

Kehao Zhang, Shaolei Zhang, Yang Feng

Model merging has emerged as an efficient strategy for constructing multitask models by integrating the strengths of multiple available expert models, thereby reducing the need to…

cs.CL2025

StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model

Shoutao Guo, Xiang Li, Mengge Liu +2

Streaming speech translation (StreamST) requires determining appropriate timing, known as policy, to generate translations while continuously receiving source speech inputs, balanc…

cs.CL2025

FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing

Shoutao Guo, Shaolei Zhang, Qingkai Fang +3

The rapid advancement of Large Language Models (LLMs) has spurred significant progress in Large Speech-Language Models (LSLMs), enhancing their capabilities in both speech understa…

cs.AI20251 cited

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Shaolei Zhang, Shoutao Guo, Qingkai Fang +2

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal intera…