13 papers
PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
Shuai Yang, Bingjie Gao, Ziwei Liu +3
Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time a…
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
Mengchen Zhang, Qi Chen, Tong Wu +2
Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-s…
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
Yuhan Zhang, Long Zhuo, Ziyang Chu +5
Despite rapid advances in 3D content generation, quality assessment for the generated 3D assets remains challenging. Existing methods mainly rely on image-based metrics and operate…
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
Yuhan Zhang, Mengchen Zhang, Tong Wu +4
3D generation is experiencing rapid advancements, while the development of 3D evaluation has not kept pace. How to keep automatic evaluation equitably aligned with human perception…
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
Jan Ackermann, Kiyohiro Nakayama, Guandao Yang +2
Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplore…
Video World Models with Long-term Spatial Memory
Tong Wu, Shuai Yang, Ryan Po +4
Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal…