37 citations · 139 across the 15 of their papers we have counts for
27 papers
Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories
Xitong Yang, Haoqi Fan, Lorenzo Torresani +2
The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We…
A Multi-View Approach To Audio-Visual Speaker Verification
Leda Sarı, Kritika Singh, Jiatong Zhou +3
Although speaker verification has conventionally been an audio-only task, some practical applications provide both audio and visual streams of input. In these cases, the visual str…
Is Space-Time Attention All You Need for Video Understanding?
Gedas Bertasius, Heng Wang, Lorenzo Torresani
We present a convolution-free approach to video classification built exclusively on self-attention over space and time. Our method, named "TimeSformer," adapts the standard Transfo…
VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
Xudong Lin, Gedas Bertasius, Jue Wang +3
We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, whic…
Resolution-Based Distillation for Efficient Histology Image Classification
Joseph DiPalma, Arief A. Suriawinata, Laura J. Tafe +2
Developing deep learning models to analyze histology images has been computationally challenging, as the massive size of the images causes excessive strain on all parts of the comp…
A Petri Dish for Histopathology Image Analysis
Jerry Wei, Arief Suriawinata, Bing Ren +9
With the rise of deep learning, there has been increased interest in using neural networks for histopathology image analysis, a field that investigates the properties of biopsy or…