activity
20242026
collaborators

7 papers

cs.CV2026

Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

Yiqian Liu, Iuliia Kotseruba, John K. Tsotsos

In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model per…

cs.CV2025

SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks

Iuliia Kotseruba, John K. Tsotsos

Generalization of deep-learning-based (DL) computer vision algorithms to various image perturbations is hard to establish and remains an active area of research. The majority of pa…

cs.CV2024

Do Saliency Models Detect Odd-One-Out Targets? New Datasets and Evaluations

Iuliia Kotseruba, Calden Wloka, Amir Rasouli +1

Recent advances in the field of saliency have concentrated on fixation prediction, with benchmarks reaching saturation. However, there is an extensive body of works in psychology a…

cs.CV2024

Diving Deeper Into Pedestrian Behavior Understanding: Intention Estimation, Action Prediction, and Event Risk Assessment

Amir Rasouli, Iuliia Kotseruba

In this paper, we delve into the pedestrian behavior understanding problem from the perspective of three different tasks: intention estimation, action prediction, and event risk as…

cs.CV2024

SCOUT+: Towards Practical Task-Driven Drivers' Gaze Prediction

Iuliia Kotseruba, John K. Tsotsos

Accurate prediction of drivers' gaze is an important component of vision-based driver monitoring and assistive systems. Of particular interest are safety-critical episodes, such as…

cs.CV2024

Data Limitations for Modeling Top-Down Effects on Drivers' Attention

Iuliia Kotseruba, John K. Tsotsos

Driving is a visuomotor task, i.e., there is a connection between what drivers see and what they do. While some models of drivers' gaze account for top-down effects of drivers' act…