4 papers · 1 filter
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
Zhiyuan Zhang, Can Wang, Dongdong Chen +1
We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes…
I2V3D: Controllable image-to-video generation with 3D guidance
Zhiyuan Zhang, Dongdong Chen, Jing Liao
We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced gene…
SGEdit: Bridging LLM with Text2Image Generative Model for Scene Graph-based Image Editing
Zhiyuan Zhang, DongDong Chen, Jing Liao
Scene graphs offer a structured, hierarchical representation of images, with nodes and edges symbolizing objects and the relationships among them. It can serve as a natural interfa…
Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM
Can Wang, Hongliang Zhong, Menglei Chai +3
Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), rece…