4 papers
ViDscribe: Multimodal AI for Customizing Audio Description and Question Answering in Online Videos
Maryam Cheema, Sina Elahimanesh, Pooyan Fazli +1
Advances in multimodal large language models enable automatic video narration and question answering (VQA), offering scalable alternatives to labor-intensive, human-authored audio…
DescribePro: Collaborative Audio Description with Human-AI Interaction
Maryam Cheema, Sina Elahimanesh, Samuel Martin +2
Audio description (AD) makes video content accessible to millions of blind and low vision (BLV) users. However, creating high-quality AD involves a trade-off between the precision…
Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
Maryam Cheema, Hasti Seifi, Pooyan Fazli
Audio descriptions (AD) make videos accessible for blind and low vision (BLV) users by describing visual elements that cannot be understood from the main audio track. AD created by…
VideoA11y: Method and Dataset for Accessible Video Description
Chaoyu Li, Sid Padmanabhuni, Maryam Cheema +2
Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall…