4 papers
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
Jesse Atuhurra, Iqra Ali, Tomoya Iwakura +2
Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evaluation of VLMs remains limited to predominantly English-centric benchmarks in whi…
J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe +1
We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scen…
NERsocial: Efficient Named Entity Recognition Dataset Construction for Human-Robot Interaction Utilizing RapidNER
Jesse Atuhurra, Hidetaka Kamigaito, Hiroki Ouchi +2
Adapting named entity recognition (NER) methods to new domains poses significant challenges. We introduce RapidNER, a framework designed for the rapid deployment of NER systems thr…
Leveraging Large Language Models in Human-Robot Interaction: A Critical Analysis of Potential and Pitfalls
Jesse Atuhurra
The emergence of large language models (LLM) and, consequently, vision language models (VLM) has ignited new imaginations among robotics researchers. At this point, the range of ap…