activity
20162022
most citedSocratic Models: Composing Zero-Shot Multimodal Reasoning with Language

173 citations · 234 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

27 papers · 1 filter

cs.CV2022173 cited

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Andy Zeng, Maria Attarian, Brian Ichter +10

Large pretrained (e.g., "foundation") models exhibit distinct capabilities depending on the domain of data they are trained on. While these domains are generic, they may only barel…

cs.CV2021

4D-Net for Learned Multi-Modal Alignment

AJ Piergiovanni, Vincent Casser, Michael S. Ryoo +1

We present 4D-Net, a 3D object detection approach, which utilizes 3D Point Cloud and RGB sensing information, both in time. We are able to incorporate the 4D information by perform…

cs.CV20213 cited

Unsupervised Discovery of Actions in Instructional Videos

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos. Instructional videos contain complex activities a…

cs.CV20214 cited

Unsupervised Action Segmentation for Instructional Videos

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. W…

cs.CV2021

Adaptive Intermediate Representations for Video Understanding

Juhana Kangaspunta, AJ Piergiovanni, Rico Jonschkowski +2

A common strategy to video understanding is to incorporate spatial and motion information by fusing features derived from RGB frames and optical flow. In this work, we introduce a…

cs.CV2021

Coarse-Fine Networks for Temporal Activity Detection in Videos

Kumara Kahatapitiya, Michael S. Ryoo

In this paper, we introduce Coarse-Fine Networks, a two-stream architecture which benefits from different abstractions of temporal resolution to learn better video representations…