1 citations · 2 across the 3 of their papers we have counts for
4 papers
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
Hongxi Li, Tong Wang, Chengjing Wu +6
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background in…
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
Zijie Lou, Xiangwei Feng, Jiaxin Wang +7
Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This app…
PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
Henghui Ding, Chang Liu, Yunchao Wei +34
Pixel-level Video Understanding in the Wild Challenge (PVUW) focus on complex video understanding. In this CVPR 2024 workshop, we add two new tracks, Complex Video Object Segmentat…
2nd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
Zhensong Xu, Jiangtao Yao, Chengjing Wu +2
Complex video object segmentation serves as a fundamental task for a wide range of downstream applications such as video editing and automatic data annotation. Here we present the…