most citedSpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

5 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20241 cited

UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model

Zhaowei Li, Wei Wang, YiQing Cai +7

Significant advancements has recently been achieved in the field of multi-modal large language models (MLLMs), demonstrating their remarkable capabilities in understanding and reas…

cs.CL2024

SpeechAlign: Aligning Speech Generation to Human Preferences

Dong Zhang, Zhaowei Li, Shimin Li +4

Speech language models have significantly advanced in generating realistic speech, with neural codec language models standing out. However, the integration of human feedback to ali…

cs.CL2024

InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Pengyu Wang, Dong Zhang, Linyang Li +5

With the rapid development of large language models (LLMs), they are not only used as general-purpose AI assistants but are also customized through further fine-tuning to meet the…

cs.CL20241 cited

SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems

Dong Zhang, Zhaowei Li, Pengyu Wang +3

Human communication is a complex and diverse process that not only involves multiple factors such as language, commonsense, and cultural backgrounds but also requires the participa…

cs.CL20235 cited

SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Dong Zhang, Shimin Li, Xin Zhang +4

Multi-modal large language models are regarded as a crucial step towards Artificial General Intelligence (AGI) and have garnered significant interest with the emergence of ChatGPT.…