5 papers · 1 filter
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
Bo Fang, Xinyao Zhang, Yuxin Song +3
Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on promptin…
CogRail: Benchmarking VLMs in Cognitive Intrusion Perception for Intelligent Railway Transportation Systems
Yonglin Tian, Qiyao Zhang, Wei Xu +9
Accurate and early perception of potential intrusion targets is essential for ensuring the safety of railway transportation systems. However, most existing systems focus narrowly o…
DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design
Xiwei Hu, Haokun Chen, Zhongqi Qi +4
We present DreamPoster, a Text-to-Image generation framework that intelligently synthesizes high-quality posters from user-provided images and text prompts while maintaining conten…
CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
Dexiang Hong, Zhao Zhang, Weidong Chen +5
Graphic design plays a crucial role in both commercial and personal contexts, yet creating high-quality, editable, and aesthetically pleasing graphic compositions remains a time-co…
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
Zhen Li, Duan Li, Yukai Guo +9
Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and str…