1 paper
Soumya Shamarao Jahagirdar, Jayasree Saha, C V Jawahar
Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging whe…