3 papers
cs.CV2025
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
Bahram Mohammadi, Ehsan Abbasnejad, Yuankai Qi +3
The remote embodied referring expression (REVERIE) task requires an agent to navigate through complex indoor environments and localize a remote object specified by high-level instr…
cs.CV2025
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
Ta Duc Huy, Duy Anh Huynh, Yutong Xie +10
Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability…
cs.CV2025
Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models
Ankit Yadav, Lingqiao Liu, Yuankai Qi
This work investigates the capabilities of current vision-language models (VLMs) in visual understanding and attribute measurement of primitive shapes using a benchmark focused on…