6 papers
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1
Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require, and finding the right frames d…
Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo +1
We introduce the Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an a…
Dress&Dance: Dress up and Dance as You Like It - Technical Preview
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo +1
We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments…
SceneCraft: Layout-Guided 3D Scene Generation
Xiuyu Yang, Yunze Man, Jun-Kun Chen +1
The creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have…
V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes
Yanming Zhang, Jun-Kun Chen, Jipeng Lyu +1
This paper introduces VEdit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content…
ProEdit: Simple Progression is All You Need for High-Quality 3D Scene Editing
Jun-Kun Chen, Yu-Xiong Wang
This paper proposes ProEdit - a simple yet effective framework for high-quality 3D scene editing guided by diffusion distillation in a novel progressive manner. Inspired by the cru…