10 papers
ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI
Hong Yang, Basura Fernando
Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and com…
Improving Temporal Action Segmentation via Constraint-Aware Decoding
Yeo Keat Ee, Debaditya Roy, Chen Li +2
Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challenges such as action variabil…
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
Chinthani Sugandhika, Chen Li, Deepu Rajan +1
Large Video-Language Models (Video-LMs) have achieved impressive progress in multimodal understanding, yet their reasoning remains weakly grounded in space and time. We present Kno…
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
Basura Fernando, Thanh-Son Nguyen, Hong Yang +3
In this work we present Knowledge Module Learning (KML) to understand and reason over procedural tasks that requires models to learn structured and compositional procedural knowled…
Explicit World Models for Reliable Human-Robot Collaboration
Kenneth Kwok, Basura Fernando, Qianli Xu +3
This paper addresses the topic of robustness under sensing noise, ambiguous instructions, and human-robot interaction. We take a radically different tack to the issue of reliable e…
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
Chinthani Sugandhika, Chen Li, Deepu Rajan +1
Spatio-temporal scene graph generation (ST-SGG) aims to model objects and their evolving relationships across video frames, enabling interpretable representations for downstream re…