activity
20182026
most citedWeakly-Supervised Online Action Segmentation in Multi-View Instructional Videos

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding

Chenyi Kuang, Nakul Agarwal

Human object interaction (HOI), gaze pattern, and their anticipation are intricately linked, providing valuable insights into cognitive processes, intentions, and behavior. However…

cs.CV2026

Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes

Nakul Agarwal, Yi-Ting Chen, Behzad Dariush

Achieving zero-collision mobility remains a key objective for intelligent vehicle systems, which requires understanding driver risk perception-a complex cognitive process shaped by…

cs.CV2025

Pose-Aware Weakly-Supervised Action Segmentation

Seth Z. Zhao, Reza Ghoddoosian, Isht Dwivedi +2

Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately l…

cs.CV2024

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos

Reza Ghoddoosian, Nakul Agarwal, Isht Dwivedi +1

Vision-language models (VLMs) are capable of recognizing unseen actions. However, existing VLMs lack intrinsic understanding of procedural action concepts. Hence, they overfit to f…

cs.CV2024

M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models

Seunggeun Chi, Hyung-gun Chi, Hengbo Ma +4

We introduce the Multi-Motion Discrete Diffusion Models (M2D2M), a novel approach for human motion generation from textual descriptions of multiple actions, utilizing the strengths…

cs.CV2024

Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models

Himangi Mittal, Nakul Agarwal, Shao-Yuan Lo +1

We introduce PlausiVL, a large video-language model for anticipating action sequences that are plausible in the real-world. While significant efforts have been made towards anticip…