collaborators

6 papers

cs.CV2026

Orient Anything V2: Unifying Orientation and Rotation Understanding

Zehan Wang, Ziang Zhang, Jiayang Xu +5

This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orie…

eess.AS2025

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang +4

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like p…

eess.AS2025

MEDIC: Zero-shot Music Editing with Disentangled Inversion Control

Huadai Liu, Jialei Wang, Xiangtai Li +6

Text-guided diffusion models revolutionize audio generation by adapting source audio to specific text prompts. However, existing zero-shot audio editing methods such as DDIM invers…

eess.AS2025

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

Yu Zhang, Changhao Pan, Wenxiang Guo +15

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datas…

eess.AS2025

FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

Huadai Liu, Jialei Wang, Rongjie Huang +4

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational deman…

cs.CV2025

Depth Anything with Any Prior

Zehan Wang, Siyu Chen, Lihe Yang +4

This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in…