10 papers · 1 filter
ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics
Zanwei Zhou, Jiazhong Cen, Jiemin Fang +7
Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plausible, fine-grained motion a…
UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement
Jingwei Yang, Ruoxi Wu, Wei Shen +4
Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, leading to reference-content l…
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
Yuhai Deng, Huimin She, Wei Shen +4
Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-quality large-scale triplet dat…
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
Yushuai Song, Weize Quan, Weining Wang +8
Recent advances in generative super-resolution (SR) have greatly improved visual realism, yet existing evaluation and optimization frameworks remain misaligned with human perceptio…
Text-Image Conditioned 3D Generation
Jiazhong Cen, Jiemin Fang, Sikuang Li +8
High-quality 3D assets are essential for VR/AR, industrial design, and entertainment, motivating growing interest in generative models that create 3D content from user prompts. Mos…
WorldGrow: Generating Infinite 3D World
Sikuang Li, Chen Yang, Jiemin Fang +6
We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face ke…