activity
20242026
collaborators

13 papers

cs.CV2026

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

Shuai Yang, Bingjie Gao, Ziwei Liu +3

Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time a…

cs.CV2025

ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation

Mengchen Zhang, Qi Chen, Tong Wu +2

Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-s…

cs.CV2025

Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity

Yuhan Zhang, Long Zhuo, Ziyang Chu +5

Despite rapid advances in 3D content generation, quality assessment for the generated 3D assets remains challenging. Existing methods mainly rely on image-based metrics and operate…

cs.CV2025

3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models

Yuhan Zhang, Mengchen Zhang, Tong Wu +4

3D generation is experiencing rapid advancements, while the development of 3D evaluation has not kept pace. How to keep automatic evaluation equitably aligned with human perception…

cs.CV2025

Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation

Jan Ackermann, Kiyohiro Nakayama, Guandao Yang +2

Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplore…

cs.CV2025

Video World Models with Long-term Spatial Memory

Tong Wu, Shuai Yang, Ryan Po +4

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal…