26 citations · 28 across the 2 of their papers we have counts for
4 papers
End-to-end Tracking with a Multi-query Transformer
Bruno Korbar, Andrew Zisserman
Multiple-object tracking (MOT) is a challenging task that requires simultaneous reasoning about location, appearance, and identity of the objects in the scene over time. Our aim in…
Video Understanding as Machine Translation
Bruno Korbar, Fabio Petroni, Rohit Girdhar +1
With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of vi…
SCSampler: Sampling Salient Clips from Video for Efficient Action Recognition
Bruno Korbar, Du Tran, Lorenzo Torresani
While many action recognition datasets consist of collections of brief, trimmed videos each containing a relevant action, videos in the real-world (e.g., on YouTube) exhibit very d…
Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization
Bruno Korbar, Du Tran, Lorenzo Torresani
There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio an…