papers

Publications (9)

cs.SD2021

Multimodal Self-Supervised Learning of General Audio Representations

Luyu Wang, Pauline Luc, Adria Recasens +2

We present a multimodal framework to learn general audio representations from videos. Existing contrastive audio representation learning methods mainly focus on using the audio mod…

cs.AI2020

Game Plan: What AI can do for Football, and What Football can do for AI

Karl Tuyls, Shayegan Omidshafiei, Paul Muller +33

The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball,…

cs.SD2022

Towards Learning Universal Audio Representations

Luyu Wang, Pauline Luc, Yan Wu +8

The ability to learn universal audio representations that can solve diverse speech, music, and environment tasks can spur many applications that require general sound content under…

cs.CV2017

Understanding Infographics through Textual and Visual Tag Prediction

Zoya Bylinskii, Sami Alsheikh, Spandan Madan +5

We introduce the problem of visual hashtag discovery for infographics: extracting visual elements from an infographic that are diagnostic of its topic. Given an infographic as inpu…

astro-ph.GA2021

A Deep Learning Approach for Characterizing Major Galaxy Mergers

Skanda Koppula, Victor Bapst, Marc Huertas-Company +15

Fine-grained estimation of galaxy merger stages from observations is a key problem useful for validation of our current theoretical understanding of galaxy formation. To this end,…

cs.CV2020

Context Based Emotion Recognition using EMOTIC Dataset

Ronak Kosti, Jose M. Alvarez, Adria Recasens +1

In our everyday lives and social interactions we often try to perceive the emotional states of people. There has been a lot of research in providing machines with a similar capacit…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.CV2022

HiP: Hierarchical Perceiver

Joao Carreira, Skanda Koppula, Daniel Zoran +10

General perception systems such as Perceivers can process arbitrary modalities in any combination and are able to handle up to a few hundred thousand inputs. They achieve this gene…

cs.CV2019

Gaze360: Physically Unconstrained Gaze Estimation in the Wild

Petr Kellnhofer, Adria Recasens, Simon Stent +2

Understanding where people are looking is an informative social cue. In this work, we present Gaze360, a large-scale gaze-tracking dataset and method for robust 3D gaze estimation…