Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
Hongxi Li, Tong Wang, Chengjing Wu +6
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background in…
cs.CV2024
Video Summarization using Denoising Diffusion Probabilistic Model
Zirui Shang, Yubo Zhu, Hongxi Li +2
Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative…
cs.CV2024
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
Yayun Qi, Hongxi Li, Yiqi Song +2
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelli…