3 papers
cs.MM2025
MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan +2
Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, b…
cs.LG2025
Coreset Selection via LLM-based Concept Bottlenecks
Akshay Mehra, Trisha Mittal, Subhadra Gopalakrishnan +1
Coreset Selection (CS) aims to identify a subset of the training dataset that achieves model performance comparable to using the entire dataset. Many state-of-the-art CS methods se…
cs.CV2025
V-Trans4Style: Visual Transition Recommendation for Video Production Style Adaptation
Pooja Guhan, Tsung-Wei Huang, Guan-Ming Su +2
We introduce V-Trans4Style, an innovative algorithm tailored for dynamic video content editing needs. It is designed to adapt videos to different production styles like documentari…