activity
20242026
collaborators

8 papers

cs.RO2026

MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction

Joerg Deigmoeller, Nakul Agarwal, Stephan Hasler +8

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires c…

cs.CV2026

Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes

Nakul Agarwal, Yi-Ting Chen, Behzad Dariush

Achieving zero-collision mobility remains a key objective for intelligent vehicle systems, which requires understanding driver risk perception-a complex cognitive process shaped by…

cs.CV2025

Task-Aware Resolution Optimization for Visual Large Language Models

Weiqing Luo, Zhen Tan, Yifan Li +4

Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fi…

cs.RO2025

CARMA: Context-Aware Situational Grounding of Human-Robot Group Interactions by Combining Vision-Language Models with Object and Action Recognition

Joerg Deigmoeller, Stephan Hasler, Nakul Agarwal +8

We introduce CARMA, a system for situational grounding in human-robot group interactions. Effective collaboration in such group settings requires situational awareness based on a c…

cs.AI2025

Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence Challenge

Weihua Du, Qiushi Lyu, Jiaming Shan +8

We introduce Constrained Human-AI Cooperation (CHAIC), an inclusive embodied social intelligence challenge designed to test social perception and cooperation in embodied agents. In…

cs.CV2025

Pose-Aware Weakly-Supervised Action Segmentation

Seth Z. Zhao, Reza Ghoddoosian, Isht Dwivedi +2

Understanding human behavior is an important problem in the pursuit of visual intelligence. A challenge in this endeavor is the extensive and costly effort required to accurately l…