53 citations · 96 across the 6 of their papers we have counts for
6 papers
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini +5
We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namel…
Augmenting Convolutional networks with attention-based aggregation
Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby +4
We show how to augment any convolutional network with an attention-based global map to achieve non-local reasoning. We replace the final average pooling by an attention-based aggre…
Interferences in match kernels
Naila Murray, Hervé Jégou, Florent Perronnin +1
We consider the design of an image representation that embeds and aggregates a set of local descriptors into a single vector. Popular representations of this kind include the bag-o…
Approximate search with quantized sparse representations
Himalaya Jain, Patrick Pérez, Rémi Gribonval +2
This paper tackles the task of storing a large collection of vectors, such as visual descriptors, and of searching in it. To this end, we propose to approximate database vectors by…
A comparison of dense region detectors for image search and fine-grained classification
Ahmet Iscen, Giorgos Tolias, Philippe-Henri Gosselin +1
We consider a pipeline for image classification or search based on coding approaches like Bag of Words or Fisher vectors. In this context, the most common approach is to extract th…
Balancing clusters to reduce response time variability in large scale image search
Romain Tavenard, Laurent Amsaleg, Hervé Jégou
Many algorithms for approximate nearest neighbor search in high-dimensional spaces partition the data into clusters. At query time, in order to avoid exhaustive search, an index se…