activity
20242026
collaborators

6 papers

cs.CV2026

Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1

Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require, and finding the right frames d…

cs.CV2025

Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview

Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo +1

We introduce the Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an a…

cs.CV2025

Dress&Dance: Dress up and Dance as You Like It - Technical Preview

Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo +1

We present Dress&Dance, a video diffusion framework that generates high quality 5-second-long 24 FPS virtual try-on videos at 1152x720 resolution of a user wearing desired garments…

cs.CV2025

SceneCraft: Layout-Guided 3D Scene Generation

Xiuyu Yang, Yunze Man, Jun-Kun Chen +1

The creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have…

cs.CV2025

V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes

Yanming Zhang, Jun-Kun Chen, Jipeng Lyu +1

This paper introduces VEdit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content…

cs.CV2024

ProEdit: Simple Progression is All You Need for High-Quality 3D Scene Editing

Jun-Kun Chen, Yu-Xiong Wang

This paper proposes ProEdit - a simple yet effective framework for high-quality 3D scene editing guided by diffusion distillation in a novel progressive manner. Inspired by the cru…