activity
20242026
collaborators

9 papers

cs.CV2026

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

Xianghao Zang, Zijian Jiang, Jiarong Cheng +8

Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…

cs.CV2026

SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls

Qianxun Xu, Chenxi Song, Yujun Cai +1

Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent videos synthesis. However, current models are predominantly optimized for singl…

cs.AI2026

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

Xiang Li, Jiabao Gao, Sipei Lin +5

The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether th…

cs.CL2026

DRIFT: Detecting Representational Inconsistencies for Factual Truthfulness

Rohan Bhatnagar, Youran Sun, Chi Andrew Zhang +2

LLMs often produce fluent but incorrect answers, yet detecting such hallucinations typically requires multiple sampling passes or post-hoc verification, adding significant latency…

cs.CV2025

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

Xiaoxing You, Qiang Huang, Lingyu Li +4

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances,…

cs.CV2025

Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness

Beier Zhu, Jiequan Cui, Hanwang Zhang +1

While image-text foundation models have succeeded across diverse downstream tasks, they still face challenges in the presence of spurious correlations between the input and label.…