8 citations · 36 across the 21 of their papers we have counts for
8 papers · 1 filter
GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
Kei Katsumata, Yui Iioka, Naoki Hosomi +3
We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging bec…
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
Kaiwen Zhou, Kwonjoon Lee, Teruhisa Misu +1
In our work, we explore the synergistic capabilities of pre-trained vision-and-language models (VLMs) and large language models (LLMs) on visual commonsense reasoning (VCR) problem…
Driving Anomaly Detection Using Conditional Generative Adversarial Network
Yuning Qiu, Teruhisa Misu, Carlos Busso
Anomaly driving detection is an important problem in advanced driver assistance systems (ADAS). It is important to identify potential hazard scenarios as early as possible to avoid…
Grounding Human-to-Vehicle Advice for Self-driving Vehicles
Jinkyu Kim, Teruhisa Misu, Yi-Ting Chen +2
Recent success suggests that deep neural control networks are likely to be a key component of self-driving vehicles. These networks are trained on large datasets to imitate human a…
Unsupervised Data Uncertainty Learning in Visual Retrieval Systems
Ahmed Taha, Yi-Ting Chen, Teruhisa Misu +2
We introduce an unsupervised formulation to estimate heteroscedastic uncertainty in retrieval systems. We propose an extension to triplet loss that models data uncertainty for each…
Exploring Uncertainty in Conditional Multi-Modal Retrieval Systems
Ahmed Taha, Yi-Ting Chen, Xitong Yang +2
We cast visual retrieval as a regression problem by posing triplet loss as a regression loss. This enables epistemic uncertainty estimation using dropout as a Bayesian approximatio…