5 papers
ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency
Shih-Po Lee, Reza Ghoddoosian, Faizan Siddiqui +2
An efficient and accurate system for detecting errors in procedural tasks is crucial for supporting human needs in daily life, as it can provide instant notifications and guide peo…
MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction
Joerg Deigmoeller, Nakul Agarwal, Stephan Hasler +8
We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires c…
CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition
Joerg Deigmoeller, Stephan Hasler, Nakul Agarwal +8
We introduce CARMA, a system for situational grounding in human-robot group interactions. Effective collaboration in such group settings requires situational awareness based on a c…
Pose-Aware Weakly-Supervised Action Segmentation
Seth Z. Zhao, Reza Ghoddoosian, Isht Dwivedi +2
Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately l…
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
Reza Ghoddoosian, Nakul Agarwal, Isht Dwivedi +1
Vision-language models (VLMs) are capable of recognizing unseen actions. However, existing VLMs lack intrinsic understanding of procedural action concepts. Hence, they overfit to f…