10 citations · 23 across the 5 of their papers we have counts for
9 papers
YORO -- Lightweight End to End Visual Grounding
Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani +2
We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object referred via natural…
Towards Differential Relational Privacy and its use in Question Answering
Simone Bombari, Alessandro Achille, Zijian Wang +6
Memorization of the relation between entities in a dataset can lead to privacy issues when using a trained model for question answering. We introduce Relational Memorization (RM) t…
DocFormer: End-to-End Transformer for Document Understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota +2
We present DocFormer -- a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand docu…
Towards Good Practices in Self-supervised Representation Learning
Srikar Appalaraju, Yi Zhu, Yusheng Xie +1
Self-supervised representation learning has seen remarkable progress in the last few years. More recently, contrastive instance learning has shown impressive results compared to it…
Saliency Driven Perceptual Image Compression
Yash Patel, Srikar Appalaraju, R. Manmatha
This paper proposes a new end-to-end trainable model for lossy image compression, which includes several novel components. The method incorporates 1) an adequate perceptual similar…
Unbiased Evaluation of Deep Metric Learning Algorithms
Istvan Fehervari, Avinash Ravichandran, Srikar Appalaraju
Deep metric learning (DML) is a popular approach for images retrieval, solving verification (same or not) problems and addressing open set classification. Arguably, the most common…