6 papers
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
Dung Nguyen, Minh Khoi Ho, Huy Ta +11
Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradic…
SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation
Vishal Thengane, Zhaochong An, Tianjin Huang +5
Incremental Few-Shot (IFS) segmentation aims to learn new categories over time from only a few annotations. Although widely studied in 2D, it remains underexplored for 3D point clo…
GDA-YOLO11: Amodal Instance Segmentation for Occlusion-Robust Robotic Fruit Harvesting
Caner Beldek, Emre Sariyildiz, Son Lam Phung +1
Occlusion remains a critical challenge in robotic fruit harvesting, as undetected or inaccurately localised fruits often results in substantial crop losses. To mitigate this issue,…
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
Ta Duc Huy, Duy Anh Huynh, Yutong Xie +10
Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability…
Multi-vision-based Picking Point Localisation of Target Fruit for Harvesting Robots
C. Beldek, A. Dunn, J. Cunningham +3
This paper presents multi-vision-based localisation strategies for harvesting robots. Identifying picking points accurately is essential for robotic harvesting because insecure gra…
Sensing-based Robustness Challenges in Agricultural Robotic Harvesting
C. Beldek, J. Cunningham, M. Aydin +3
This paper presents the challenges agricultural robotic harvesters face in detecting and localising fruits under various environmental disturbances. In controlled laboratory settin…