7 papers
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos
In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model per…
SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
Iuliia Kotseruba, John K. Tsotsos
Generalization of deep-learning-based (DL) computer vision algorithms to various image perturbations is hard to establish and remains an active area of research. The majority of pa…
Do Saliency Models Detect Odd-One-Out Targets? New Datasets and Evaluations
Iuliia Kotseruba, Calden Wloka, Amir Rasouli +1
Recent advances in the field of saliency have concentrated on fixation prediction, with benchmarks reaching saturation. However, there is an extensive body of works in psychology a…
Diving Deeper Into Pedestrian Behavior Understanding: Intention Estimation, Action Prediction, and Event Risk Assessment
Amir Rasouli, Iuliia Kotseruba
In this paper, we delve into the pedestrian behavior understanding problem from the perspective of three different tasks: intention estimation, action prediction, and event risk as…
SCOUT+: Towards Practical Task-Driven Drivers' Gaze Prediction
Iuliia Kotseruba, John K. Tsotsos
Accurate prediction of drivers' gaze is an important component of vision-based driver monitoring and assistive systems. Of particular interest are safety-critical episodes, such as…
Data Limitations for Modeling Top-Down Effects on Drivers' Attention
Iuliia Kotseruba, John K. Tsotsos
Driving is a visuomotor task, i.e., there is a connection between what drivers see and what they do. While some models of drivers' gaze account for top-down effects of drivers' act…