26 papers
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
Renlong Wu, Haoran Chen, Yuxiang Wei +3
The paper introduces 4DHumanDiff, a diffusion-based framework that directly generates 360-degree dynamic human models as 4D Gaussian Splatting representations from text prompts, el…
InstanceControl: Controllable Complex Image Generation without Instance Labeling
Xiaoyu Liu, Huan Wang, Fan Li +4
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. Howev…
Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer
Bo Wei, Xianhui Lin, Yi Dong +8
Makeup transfer applies a reference cosmetic style to a source face while preserving its identity and geometry. However, this task is severely hindered by the lack of real paired t…
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Chonghuinan Wang, Zhikai Chen, Chunwei Wang +9
The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving…
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
Haoze Zhang, Tianyu Huang, Zichen Wan +4
While recent video generation models have achieved significant visual fidelity, they often suffer from the lack of explicit physical controllability and plausibility. To address th…
Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning
Mengzhao Wang, Yanli Ji, Wangmeng Zuo +2
Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeated…