collaborators

6 papers

cs.CV2026

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

Dongxu Ge, Shansong Liu, Cheng Gong +3

As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted incre…

eess.AS2026

Unified Audio Generation and Editing via Joint Condition Modeling and Progressive Training

Haocheng Dong, Yuheng Lu, Cheng Gong +3

With the growing focus on audio in multimedia applications, numerous advanced works on audio generation have emerged. Existing studies typically treat text-to-audio (TTA) and other…

eess.AS2026

High-Fidelity Generative Audio Compression at 0.275kbps

Hao Ma, Ruihao Jing, Shansong Liu +4

High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditiona…

eess.AS2025

Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models

Ruihao Jing, Cheng Gong, Yu Jiang +5

Rare words remain a critical bottleneck for speech-to-text systems. While direct fine-tuning improves recognition of target words, it often incurs high cost, catastrophic forgettin…

eess.AS2025

Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization

Xueqing Li, Hao Ma, Zehan Li +8

Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…

eess.AS2025

: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation

Boyu Zhu, Cheng Gong, Muyang Wu +5

Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these mode…