collaborators

6 papers

cs.RO2026

Explicit World Models for Reliable Human-Robot Collaboration

Kenneth Kwok, Basura Fernando, Qianli Xu +3

This paper addresses the topic of robustness under sensing noise, ambiguous instructions, and human-robot interaction. We take a radically different tack to the issue of reliable e…

cs.CV2025

VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation

Chinthani Sugandhika, Chen Li, Deepu Rajan +1

Spatio-temporal scene graph generation (ST-SGG) aims to model objects and their evolving relationships across video frames, enabling interpretable representations for downstream re…

cs.CV2025

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

Chen Li, Eric Peh, Basura Fernando

Recent advances in large vision-language models (VLMs) have shown significant promise for 3D scene understanding. Existing VLM-based approaches typically align 3D scene features wi…

cs.CV2025

Mitigating Easy Option Bias in Multiple-Choice Question Answering

Hao Zhang, Chen Li, Basura Fernando

In this early study, we observe an Easy-Options Bias (EOB) issue in some multiple-choice Visual Question Answering (VQA) benchmarks such as MMStar, RealWorldQA, SEED-Bench, Next-QA…

cs.CV2025

IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A

Chen Li, Chinthani Sugandhika, Yeo Keat Ee +5

Existing human motion Q\&A methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To…

cs.CV2025

Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering

Basura Fernando, Thanh-Son Nguyen, Hong Yang +3

In this work we present Knowledge Module Learning (KML) to understand and reason over procedural tasks that requires models to learn structured and compositional procedural knowled…