4 papers
LACON: Training Text-to-Image Model from Uncurated Data
Zhiyang Liang, Ziyu Wan, Hongyu Liu +4
The success of modern text-to-image generation is largely attributed to massive, high-quality datasets. Currently, these datasets are curated through a filter-first paradigm that a…
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
Zhiyuan Zhang, Can Wang, Dongdong Chen +1
We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes…
I2V3D: Controllable image-to-video generation with 3D guidance
Zhiyuan Zhang, Dongdong Chen, Jing Liao
We present I2V3D, a novel framework for animating static images into dynamic videos with precise 3D control, leveraging the strengths of both 3D geometry guidance and advanced gene…
SGEdit: Bridging LLM with Text2Image Generative Model for Scene Graph-based Image Editing
Zhiyuan Zhang, DongDong Chen, Jing Liao
Scene graphs offer a structured, hierarchical representation of images, with nodes and edges symbolizing objects and the relationships among them. It can serve as a natural interfa…