7 papers
Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control
Boyu Li, Linjie Qiu, Lin-Ping Yuan +4
Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users in the iterative refinement of generati…
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang +23
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community tr…
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
Gongye Liu, Bo Yang, Yida Zhi +8
Preference optimization for diffusion and flow-matching models relies on reward functions that are both discriminatively robust and computationally efficient. Vision-Language Model…
ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation
Hanxiao Sun, Mingxin Yang, Shuhui Yang +5
Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unseen regions. Multi-view condi…
SketchDynamics: Exploring Free-Form Sketches for Dynamic Intent Expression in Animation Generation
Boyu Li, Lin-Ping Yuan, Zeyu Wang +1
Sketching provides an intuitive way to convey dynamic intent in animation authoring (i.e., how elements change over time and space), making it a natural medium for automatic conten…
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
Yidan Zhang, Mutian Xu, Yiming Hao +6
Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and tim…