2 papers
cs.CV2026
MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
Yifan Xu, Chao Zhang, Ruifei Ma +4
The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-lev…
cs.CV2025
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
Jiawei Liu, Yuanzhi Zhu, Feiyu Gao +5
Generating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, c…