3 papers
cs.CV2026
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos
In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model per…
cs.RO2026
DIJIT: A Robotic Head for an Active Observer
Mostafa Kamali Tabrizi, Mingshi Chi, Bir Bikram Dey +5
We present DIJIT, a novel binocular robotic head expressly designed for mobile agents that behave as active observers. DIJIT's unique breadth of functionality enables active vision…
cs.LG2025
Boosting Reinforcement Learning in 3D Visuospatial Tasks Through Human-Informed Curriculum Design
Markus D. Solbach, John K. Tsotsos
Reinforcement Learning is a mature technology, often suggested as a potential route towards Artificial General Intelligence, with the ambitious goal of replicating the wide range o…