2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.CV2019★ 2 cited
Compositional Temporal Visual Grounding of Natural Language Event Descriptions
Jonathan C. Stroud, Ryan McCaffrey, Rada Mihalcea +2
Temporal grounding entails establishing a correspondence between natural language event descriptions and their visual depictions. Compositional modeling becomes central: we first g…
cs.CV2018
D3D: Distilled 3D Networks for Video Action Recognition
Jonathan C. Stroud, David A. Ross, Chen Sun +2
State-of-the-art methods for video action recognition commonly use an ensemble of two networks: the spatial stream, which takes RGB frames as input, and the temporal stream, which…