130 citations · 132 across the 3 of their papers we have counts for
3 papers
cs.MM2023
Audio Retrieval for Multimodal Design Documents: A New Dataset and Algorithms
Prachi Singh, Srikrishna Karanam, Sumit Shekhar
We consider and propose a new problem of retrieving audio files relevant to multimodal design document inputs comprising both textual elements and visual imagery, e.g., birthday/gr…
cs.CV2022★ 2 cited
PseudoClick: Interactive Image Segmentation with Click Imitation
Qin Liu, Meng Zheng, Benjamin Planche +4
The goal of click-based interactive image segmentation is to obtain precise object segmentation masks with limited user interaction, i.e., by a minimal number of user clicks. Exist…
cs.CV2021★ 130 cited
Learning Hierarchical Attention for Weakly-supervised Chest X-Ray Abnormality Localization and Diagnosis
Xi Ouyang, Srikrishna Karanam, Ziyan Wu +5
We consider the problem of abnormality localization for clinical applications. While deep learning has driven much recent progress in medical imaging, many clinical challenges are…