4 papers
CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe +4
4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited…
Explicit World Models for Reliable Human-Robot Collaboration
Kenneth Kwok, Basura Fernando, Qianli Xu +3
This paper addresses the topic of robustness under sensing noise, ambiguous instructions, and human-robot interaction. We take a radically different tack to the issue of reliable e…
On the development of an AI performance and behavioural measures for teaching and classroom management
Andreea I. Niculescu, Jochen Ehnes, Chen Yi +9
This paper presents a two-year research project focused on developing AI-driven measures to analyze classroom dynamics, with particular emphasis on teacher actions captured through…
Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju +3
3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D sce…