10 papers
Mitigating Cognitive Bias in RLHF by Altering Rationality
Tiffany Horter, Andrew Markham, Niki Trigoni +1
How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outputs are used to train a rewar…
CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
Chenyang Ma, Guangyu Yang, Kai Lu +6
Current work on robot failure detection and correction typically operates in a post hoc manner, analyzing errors and applying corrections only after failures occur. This work intro…
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
Shitong Xu, Yiyuan Yang, Niki Trigoni +1
Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's iden…
COOPERA: Continual Open-Ended Human-Robot Assistance
Chenyang Ma, Kai Lu, Ruta Desai +3
To understand and collaborate with humans, robots must account for individual human traits, habits, and activities over time. However, most robotic assistants lack these abilities,…
Data Factory with Minimal Human Effort Using VLMs
Jiaojiao Ye, Jiaxing Zhong, Qian Xie +3
Generating enough and diverse data through augmentation offers an efficient solution to the time-consuming and labour-intensive process of collecting and annotating pixel-wise imag…
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
Yiyuan Yang, Shitong Xu, Niki Trigoni +1
Sound source localization (SSL) is a critical technology for determining the position of sound sources in complex environments. However, existing methods face challenges such as hi…