8 papers · 1 filter
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
Xianghao Zang, Zijian Jiang, Jiarong Cheng +8
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…
SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
Qianxun Xu, Chenxi Song, Yujun Cai +1
Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent videos synthesis. However, current models are predominantly optimized for singl…
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
Xiaoxing You, Qiang Huang, Lingyu Li +4
News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances,…
Video-Bench: Human-Aligned Video Generation Benchmark
Hui Han, Siyuan Li, Jiaqi Chen +10
Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video g…
Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness
Beier Zhu, Jiequan Cui, Hanwang Zhang +1
While image-text foundation models have succeeded across diverse downstream tasks, they still face challenges in the presence of spurious correlations between the input and label.…
LoRA of Change: Learning to Generate LoRA for the Editing Instruction from A Single Before-After Image Pair
Xue Song, Jiequan Cui, Hanwang Zhang +4
In this paper, we propose the LoRA of Change (LoC) framework for image editing with visual instructions, i.e., before-after image pairs. Compared to the ambiguities, insufficient s…