2 papers
cs.CV2026
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
Hongxi Li, Tong Wang, Chengjing Wu +6
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background in…
cs.CV2024
Storyboard guided Alignment for Fine-grained Video Action Recognition
Enqi Liu, Liyuan Pan, Yan Yang +4
Fine-grained video action recognition can be conceptualized as a video-text matching problem. Previous approaches often rely on global video semantics to consolidate video embeddin…