activity
20162024
most citedRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

273 citations · 507 across the 16 of their papers we have counts for

collaborators
Showing 2021Show all

8 papers · 1 filter

cs.CV2021

4D-Net for Learned Multi-Modal Alignment

AJ Piergiovanni, Vincent Casser, Michael S. Ryoo +1

We present 4D-Net, a 3D object detection approach, which utilizes 3D Point Cloud and RGB sensing information, both in time. We are able to incorporate the 4D information by perform…

cs.RO20212 cited

Self-Supervised Disentangled Representation Learning for Third-Person Imitation Learning

Jinghuan Shang, Michael S. Ryoo

Humans learn to imitate by observing others. However, robot imitation learning generally requires expert demonstrations in the first-person view (FPV). Collecting such FPV videos f…

cs.CV20213 cited

Unsupervised Discovery of Actions in Instructional Videos

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos. Instructional videos contain complex activities a…

cs.CV20214 cited

Unsupervised Action Segmentation for Instructional Videos

AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo +1

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. W…

cs.CV2021

Adaptive Intermediate Representations for Video Understanding

Juhana Kangaspunta, AJ Piergiovanni, Rico Jonschkowski +2

A common strategy to video understanding is to incorporate spatial and motion information by fusing features derived from RGB frames and optical flow. In this work, we introduce a…

cs.CV2021

Coarse-Fine Networks for Temporal Activity Detection in Videos

Kumara Kahatapitiya, Michael S. Ryoo

In this paper, we introduce Coarse-Fine Networks, a two-stream architecture which benefits from different abstractions of temporal resolution to learn better video representations…