15 citations · 34 across the 5 of their papers we have counts for
4 papers · 1 filter
Using Multimodal Deep Neural Networks to Disentangle Language from Visual Aesthetics
Colin Conwell, Christopher Hamblin, Chelsea Boccagno +4
When we experience a visual stimulus as beautiful, how much of that experience derives from perceptual computations we cannot describe versus conceptual knowledge we can readily tr…
On the use of Cortical Magnification and Saccades as Biological Proxies for Data Augmentation
Binxu Wang, David Mayo, Arturo Deza +2
Self-supervised learning is a powerful way to learn useful representations from natural data. It has also been suggested as one possible means of building visual representation in…
Large-Scale Automatic Labeling of Video Events with Verbs Based on Event-Participant Interaction
Andrei Barbu, Alexander Bridge, Dan Coroian +12
We present an approach to labeling short video clips with English verbs as event descriptions. A key distinguishing aspect of this work is that it labels videos with verbs that des…
Video In Sentences Out
Andrei Barbu, Alexander Bridge, Zachary Burchill +15
We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as…