21 citations · 32 across the 3 of their papers we have counts for
4 papers
Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov +4
Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instruc…
Coarse to Fine Multi-Resolution Temporal Convolutional Network
Dipika Singhania, Rahul Rahaman, Angela Yao
Temporal convolutional networks (TCNs) are a commonly used architecture for temporal video segmentation. TCNs however, tend to suffer from over-segmentation errors and require addi…
Temporal Aggregate Representations for Long-Range Video Understanding
Fadime Sener, Dipika Singhania, Angela Yao
Future prediction, especially in long-range videos, requires reasoning from current and past observations. In this work, we address questions of temporal extent, scaling, and level…
Rethinking CNN Models for Audio Classification
Kamalesh Palanisamy, Dipika Singhania, Angela Yao
In this paper, we show that ImageNet-Pretrained standard deep CNN models can be used as strong baseline networks for audio classification. Even though there is a significant differ…