7 citations · 14 across the 3 of their papers we have counts for
5 papers
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
Md Touhidul Islam, Imran Kabir, Md Alimoor Reza +1
We present IKIWISI ("I Know It When I See It"), an interactive visual pattern generator for assessing vision-language models in video object recognition when ground truth is unavai…
Blending 3D Geometry and Machine Learning for Multi-View Stereopsis
Vibhas Vats, Md. Alimoor Reza, David Crandall +1
Traditional multi-view stereo (MVS) methods primarily depend on photometric and geometric consistency constraints. In contrast, modern learning-based algorithms often rely on the p…
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
Imran Kabir, Md Alimoor Reza, Syed Billah
Large multimodal models (LMMs) are increasingly integrated into autonomous driving systems for user interaction. However, their limitations in fine-grained spatial reasoning pose c…
Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation
Md Touhidul Islam, Imran Kabir, Elena Ariel Pearce +2
This paper presents a curated list of 90 objects essential for the navigation of blind and low-vision (BLV) individuals, encompassing road, sidewalk, and indoor environments. We de…
Active Object Manipulation Facilitates Visual Object Learning: An Egocentric Vision Study
Satoshi Tsutsui, Dian Zhi, Md Alimoor Reza +2
Inspired by the remarkable ability of the infant visual learning system, a recent study collected first-person images from children to analyze the `training data' that they receive…