Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
Bo Fang, Xinyao Zhang, Yuxin Song +3
Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on promptin…
cs.CV2024
Re-Attentional Controllable Video Diffusion Editing
Yuanzhi Wang, Yong Li, Mengyi Liu +4
Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. R…
cs.CV2024
FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models
Wei Wu, Qingnan Fan, Shuai Qin +3
Precise image editing with text-to-image models has attracted increasing interest due to their remarkable generative capabilities and user-friendly nature. However, such attempts f…