7 papers
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
Yilei Hua, Beibei Jing, Ce Zheng +3
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods…
When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation
Siwei Meng, Yawei Luo, Shu Zhang +1
Text-to-video (T2V) generation models have achieved strong visual realism, but improving physical plausibility can come at the cost of semantic consistency with the input text. Thi…
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM
Siwei Meng, Yawei Luo, Ping Liu
Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. However, state-of-the-art video diffus…
Alignment Is All You Need For X-to-4D Generation
Qiaowei Miao, Kehan Li, Yawei Luo +1
Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality-to-4D (X-to-4D) gen…
SARe: Structure-Aware Generative 3D Fragment Reassembly
Hanze Jia, Chunshi Wang, Yuxiao Yang +4
3D fragment reassembly estimates the rigid pose of each fragment to recover a complete object from unordered point clouds or meshes. The task becomes increasingly challenging as th…
Advances in 4D Generation: A Survey
Qiaowei Miao, Kehan Li, Jinsheng Quan +6
Generative artificial intelligence has recently progressed from static image and video synthesis to 3D content generation, culminating in the emergence of 4D generation-the task of…