collaborators

5 papers

cs.MM2026

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7

Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…

cs.CL2026

PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models

Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang +6

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is large…

cs.SD2025

I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception

Jiawei Zhang, Tian-Hao Zhang, Jun Wang +3

Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works hav…

cs.SD2025

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles

Tian-Hao Zhang, Jiawei Zhang, Jun Wang +2

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vis…

eess.AS2025

Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition

Wei Zhang, Tian-Hao Zhang, Chao Luo +4

Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weight…