activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization

Jiawei Mao, Xiaoke Huang, Yunfei Xie +7

This paper introduces Story-Iter, a new training-free iterative paradigm to enhance long-story generation. Unlike existing methods that rely on fixed reference images to construct…

cs.CV2025

ARFlow: Autoregressive Flow with Hybrid Linear Attention

Mude Hui, Rui-Jie Zhu, Songlin Yang +5

Flow models are effective at progressively generating realistic images, but they generally struggle to capture long-range dependencies during the generation process as they compres…

cs.CV2025

: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

Siwei Yang, Mude Hui, Bingchen Zhao +3

We introduce , a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity…

cs.CV2024

What If We Recaption Billions of Web Images with LLaMA-3?

Xianhang Li, Haoqin Tu, Mude Hui +9

Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance…

cs.CV2024

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Mude Hui, Siwei Yang, Bingchen Zhao +5

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feed…