9 papers
Context-structured Video Anomaly Detection with Large Vision-Language Models
Dongjun Kim, Changjae Oh, Andrea Cavallaro +1
Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable t…
NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation Problem
Daniel Fuertes, Andrea Cavallaro, Carlos R. del-Blanco +2
Path planning is usually solved by addressing either the (high-level) route planning problem (waypoint sequencing to achieve the final goal) or the (low-level) path planning proble…
Toward Human-Robot Teaming: Learning Handover Behaviors from 3D Scenes
Yuekun Wu, Yik Lung Pang, Andrea Cavallaro +1
Human-robot teaming (HRT) systems often rely on large-scale datasets of human and robot interactions, especially for close-proximity collaboration tasks such as human-robot handove…
Improving Generalization of Language-Conditioned Robot Manipulation
Chenglin Cui, Chaoran Zhu, Changjae Oh +1
The control of robots for manipulation tasks generally relies on visual input. Recent advances in vision-language models (VLMs) enable the use of natural language instructions to c…
Learning human-to-robot handovers through 3D scene reconstruction
Yuekun Wu, Yik Lung Pang, Andrea Cavallaro +1
Learning robot manipulation policies from raw, real-world image data requires a large number of robot-action trials in the physical environment. Although training using simulations…
Stereo Hand-Object Reconstruction for Human-to-Robot Handover
Yik Lung Pang, Alessio Xompero, Changjae Oh +1
Jointly estimating hand and object shape facilitates the grasping task in human-to-robot handovers. However, relying on hand-crafted prior knowledge about the geometric structure o…