audio generation 1diffusion models 1multimodal audio 1temporal scene synthesis 1transformer 1variational autoencoder 1
From the 1 of 3 linked papers with an AI index.
3 papers
eess.AS2026
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…
cs.CV2025
MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
Zhicheng Du, Qingyang Shi, Jiasheng Lu +4
The multimodal relevance metric is usually borrowed from the embedding ability of pretrained contrastive learning models for bimodal data, which is used to evaluate the correlation…
cs.CV2025
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
Yingshan Liang, Keyu Fan, Zhicheng Du +5
Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information strug…