1 citations · 1 across the 2 of their papers we have counts for
5 papers
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
Shravan Murlidaran, Ziqi Wen, Sana Shehabi +1
When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on people, text, objects being gaz…
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
Ziqi Wen, Parsa Madinei, Miguel P. Eckstein
Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpreta…
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
Parsa Madinei, Ryan Solgi, Ziqi Wen +3
We introduce INTERLACE, a novel framework that prunes redundant layers in VLMs while maintaining performance through sample-efficient finetuning. Existing layer pruning methods lea…
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
Jonathan Skaza, Parsa Madinei, Ziqi Wen +1
Visual complexity prediction is a fundamental problem in computer vision with applications in image compression, retrieval, and classification. Understanding what makes humans perc…
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
Ziqi Wen, Jonathan Skaza, Shravan Murlidaran +2
Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-computable predictors for scene und…