From the 1 of 12 linked papers with an AI index.
12 papers
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang +4
The paper introduces CAIRN, a topology‑aware large multimodal model that uses graph neural networks and hierarchical attention to understand and reason about multi‑room 3D scenes.
Mitigating Cognitive Bias in RLHF by Altering Rationality
Tiffany Horter, Andrew Markham, Niki Trigoni +1
How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outputs are used to train a rewar…
CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
Chenyang Ma, Guangyu Yang, Kai Lu +6
Current work on robot failure detection and correction typically operates in a post hoc manner, analyzing errors and applying corrections only after failures occur. This work intro…
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
Shitong Xu, Yiyuan Yang, Niki Trigoni +1
Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's iden…
COOPERA: Continual Open-Ended Human-Robot Assistance
Chenyang Ma, Kai Lu, Ruta Desai +3
To understand and collaborate with humans, robots must account for individual human traits, habits, and activities over time. However, most robotic assistants lack these abilities,…
Data Factory with Minimal Human Effort Using VLMs
Jiaojiao Ye, Jiaxing Zhong, Qian Xie +3
Generating enough and diverse data through augmentation offers an efficient solution to the time-consuming and labour-intensive process of collecting and annotating pixel-wise imag…