4 papers
BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation
Baoyou Chen, Hanchen Xia, Peng Tu +5
Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck. Diffusion VLMs offe…
Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition
Hao Shi, Yuan Gao, Xugang Lu +1
Large Language Models (LLMs) are effective decoders for Serialized Output Training (SOT) in two-talker automatic speech recognition (ASR), but their performance degrades substantia…
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
Yuan Gao, Hao Shi, Yahui Fu +2
This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We co…
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
Hao Shi, Xugang Lu, Kazuki Shimada +1
Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable cond…