9 papers
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
Xianghao Zang, Zijian Jiang, Jiarong Cheng +8
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…
SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
Qianxun Xu, Chenxi Song, Yujun Cai +1
Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent videos synthesis. However, current models are predominantly optimized for singl…
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
Xiang Li, Jiabao Gao, Sipei Lin +5
The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether th…
DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness
Rohan Bhatnagar, Youran Sun, Chi Andrew Zhang +2
LLMs often produce fluent but incorrect answers, yet detecting such hallucinations typically requires multiple sampling passes or post-hoc verification, adding significant latency…
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
Xiaoxing You, Qiang Huang, Lingyu Li +4
News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances,…
Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness
Beier Zhu, Jiequan Cui, Hanwang Zhang +1
While image-text foundation models have succeeded across diverse downstream tasks, they still face challenges in the presence of spurious correlations between the input and label.…