2 papers
cs.CV2024
Contextual AD Narration with Interleaved Multimodal Sequence
Hanlin Wang, Zhan Tong, Kecheng Zheng +2
The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With vi…
cs.CV2023
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
Qinying Liu, Wei Wu, Kecheng Zheng +6
The crux of learning vision-language models is to extract semantically aligned information from visual and linguistic data. Existing attempts usually face the problem of coarse ali…