2 papers
cs.CV2026
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
Bo Fang, Xinyao Zhang, Yuxin Song +3
Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on promptin…
cs.CV2024
Re-Attentional Controllable Video Diffusion Editing
Yuanzhi Wang, Yong Li, Mengyi Liu +4
Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. R…