5 citations · 6 across the 2 of their papers we have counts for
3 papers
Adaptive Intermediate Representations for Video Understanding
Juhana Kangaspunta, AJ Piergiovanni, Rico Jonschkowski +2
A common strategy to video understanding is to incorporate spatial and motion information by fusing features derived from RGB frames and optical flow. In this work, we introduce a…
AssembleNet++: Assembling Modality Representations via Attention Connections
Michael S. Ryoo, AJ Piergiovanni, Juhana Kangaspunta +1
We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy at…
SPIN: A High Speed, High Resolution Vision Dataset for Tracking and Action Recognition in Ping Pong
Steven Schwarcz, Peng Xu, David D'Ambrosio +4
We introduce a new high resolution, high frame rate stereo video dataset, which we call SPIN, for tracking and action recognition in the game of ping pong. The corpus consists of p…