4 papers · 1 filter
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
Xianghao Zang, Zijian Jiang, Jiarong Cheng +8
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Mod…
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
Runpu Wei, Zijin Yin, Shuo Zhang +8
Inpainting algorithms have achieved remarkable progress in removing objects from images, yet still face two challenges: 1) struggle to handle the object's visual effects such as sh…
Disentangle and denoise: Tackling context misalignment for video moment retrieval
Kaijing Ma, Han Fang, Xianghao Zang +7
Video Moment Retrieval, which aims to locate in-context video moments according to a natural language query, is an essential task for cross-modal grounding. Existing methods focus…
ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
Han Fang, Xianghao Zang, Chao Ban +5
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since vid…