173 citations · 234 across the 12 of their papers we have counts for
27 papers · 1 filter
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Andy Zeng, Maria Attarian, Brian Ichter +10
Large pretrained (e.g., "foundation") models exhibit distinct capabilities depending on the domain of data they are trained on. While these domains are generic, they may only barel…
4D-Net for Learned Multi-Modal Alignment
AJ Piergiovanni, Vincent Casser, Michael S. Ryoo +1
We present 4D-Net, a 3D object detection approach, which utilizes 3D Point Cloud and RGB sensing information, both in time. We are able to incorporate the 4D information by perform…
Unsupervised Discovery of Actions in Instructional Videos
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1
In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos. Instructional videos contain complex activities a…
Unsupervised Action Segmentation for Instructional Videos
AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1
In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. W…
Adaptive Intermediate Representations for Video Understanding
Juhana Kangaspunta, AJ Piergiovanni, Rico Jonschkowski +2
A common strategy to video understanding is to incorporate spatial and motion information by fusing features derived from RGB frames and optical flow. In this work, we introduce a…
Coarse-Fine Networks for Temporal Activity Detection in Videos
Kumara Kahatapitiya, Michael S. Ryoo
In this paper, we introduce Coarse-Fine Networks, a two-stream architecture which benefits from different abstractions of temporal resolution to learn better video representations…