5 papers · 1 filter
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie +7
This paper introduces Story-Iter, a new training-free iterative paradigm to enhance long-story generation. Unlike existing methods that rely on fixed reference images to construct…
ARFlow: Autoregressive Flow with Hybrid Linear Attention
Mude Hui, Rui-Jie Zhu, Songlin Yang +5
Flow models are effective at progressively generating realistic images, but they generally struggle to capture long-range dependencies during the generation process as they compres…
: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
Siwei Yang, Mude Hui, Bingchen Zhao +3
We introduce , a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity…
What If We Recaption Billions of Web Images with LLaMA-3?
Xianhang Li, Haoqin Tu, Mude Hui +9
Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance…
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Mude Hui, Siwei Yang, Bingchen Zhao +5
This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feed…