Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
Yifan Xu, Chao Zhang, Ruifei Ma +4
The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-lev…
cs.CV2025
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
Jiawei Liu, Yuanzhi Zhu, Feiyu Gao +5
Generating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, c…
cs.CV2024
XS-VID: An Extremely Small Video Object Detection Dataset
Jiahao Guo, Ziyang Xu, Lianjun Wu +3
Small Video Object Detection (SVOD) is a crucial subfield in modern computer vision, essential for early object discovery and detection. However, existing SVOD datasets are scarce…