2 papers
cs.CV2023
"Let's not Quote out of Context": Unified Vision-Language Pretraining for Context Assisted Image Captioning
Abisek Rajakumar Kalarani, Pushpak Bhattacharyya, Niyati Chhaya +1
Well-formed context aware image captions and tags in enterprise content such as marketing material are critical to ensure their brand presence and content recall. Manual creation a…
cs.MM2023
Audio Retrieval for Multimodal Design Documents: A New Dataset and Algorithms
Prachi Singh, Srikrishna Karanam, Sumit Shekhar
We consider and propose a new problem of retrieving audio files relevant to multimodal design document inputs comprising both textual elements and visual imagery, e.g., birthday/gr…