collaborators

6 papers

cs.SD2025

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters

Haomin Zhang, Kristin Qi, Shuxin Yang +3

Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned…

cs.SD2025

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

Chang Liu, Haomin Zhang, Shiyu Xia +5

Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, e…

cs.CV2025

DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Haomin Zhang, Chang Liu, Junjie Zheng +3

Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio ben…

cs.SD2025

Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization

Haomin Zhang, Sizhe Shan, Haoyu Wang +4

Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-s…

cs.SD2024

Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio

Gongyu Chen, Haomin Zhang, Chaofan Ding +2

One fascinating aspect of pre-trained Audio-Language Models (ALMs) learning is their impressive zero-shot generalization capability and test-time adaptation (TTA) methods aiming to…

cs.SD2024

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

Zihao Chen, Haomin Zhang, Xinhan Di +10

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-…