2 citations · 3 across the 6 of their papers we have counts for
10 papers · 1 filter
Cross-Modal Retrieval Augmentation for Multi-Modal Classification
Shir Gur, Natalia Neverova, Chris Stauffer +3
Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing. Here,…
Exploring Visual Engagement Signals for Representation Learning
Menglin Jia, Zuxuan Wu, Austin Reiter +3
Visual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such visual engagement clues a…
Intentonomy: a Dataset and Study towards Human Intent Understanding
Menglin Jia, Zuxuan Wu, Austin Reiter +3
An image is worth a thousand words, conveying information that goes beyond the physical visual content therein. In this paper, we study the intent behind social media images with a…
Deep Multi-Modal Sets
Austin Reiter, Menglin Jia, Pu Yang +1
Many vision-related tasks benefit from reasoning over multiple modalities to leverage complementary views of data in an attempt to learn robust embedding spaces. Most deep learning…
Action Recognition Using Volumetric Motion Representations
Michael Peven, Gregory D. Hager, Austin Reiter
Traditional action recognition models are constructed around the paradigm of 2D perspective imagery. Though sophisticated time-series models have pushed the field forward, much of…
Dense Depth Estimation in Monocular Endoscopy with Self-supervised Learning Methods
Xingtong Liu, Ayushi Sinha, Masaru Ishii +4
We present a self-supervised approach to training convolutional neural networks for dense depth estimation from monocular endoscopy data without a priori modeling of anatomy or sha…