5 papers
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
Xianghao Zang, Zijian Jiang, Jiarong Cheng +8
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…
DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents
Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia +7
Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instea…
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
Shuo Zhang, Wenzhuo Wu, Huayu Zhang +8
Recent advances in diffusion models have significantly improved image editing. However, challenges persist in handling geometric transformations, such as translation, rotation, and…
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
Canhui Tang, Zifan Han, Hongbo Sun +7
Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs.…
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
Yuxuan Wang, Tianwei Cao, Huayu Zhang +3
Image generation has achieved remarkable progress with the development of large-scale text-to-image models, especially diffusion-based models. However, generating human images with…