collaborators

8 papers

cs.CV2025

ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation

Mengchen Zhang, Qi Chen, Tong Wu +2

Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-s…

cs.CV2025

Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity

Yuhan Zhang, Long Zhuo, Ziyang Chu +5

Despite rapid advances in 3D content generation, quality assessment for the generated 3D assets remains challenging. Existing methods mainly rely on image-based metrics and operate…

cs.CV2025

Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation

Jan Ackermann, Kiyohiro Nakayama, Guandao Yang +2

Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplore…

cs.CV2025

Video World Models with Long-term Spatial Memory

Tong Wu, Shuai Yang, Ryan Po +4

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal…

cs.CV2025

GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography

Mengchen Zhang, Tong Wu, Jing Tan +3

Camera trajectory design plays a crucial role in video production, serving as a fundamental tool for conveying directorial intent and enhancing visual storytelling. In cinematograp…

cs.CV2025

3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models

Yuhan Zhang, Mengchen Zhang, Tong Wu +4

3D generation is experiencing rapid advancements, while the development of 3D evaluation has not kept pace. How to keep automatic evaluation equitably aligned with human perception…