6 papers
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
Ryan Solgi, Parsa Madinei, Jiayi Tian +4
Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment.…
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
Ziqi Wen, Parsa Madinei, Miguel P. Eckstein
Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpreta…
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
Parsa Madinei, Srijita Karmakar, Russell Cohen Hoffing +2
We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolve ambiguity in open-ended VQA. T…
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
Parsa Madinei, Ryan Solgi, Ziqi Wen +3
We introduce INTERLACE, a novel framework that prunes redundant layers in VLMs while maintaining performance through sample-efficient finetuning. Existing layer pruning methods lea…
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
Jonathan Skaza, Parsa Madinei, Ziqi Wen +1
Visual complexity prediction is a fundamental problem in computer vision with applications in image compression, retrieval, and classification. Understanding what makes humans perc…
ARChef: An iOS-Based Augmented Reality Cooking Assistant Powered by Multimodal Gemini LLM
Rithik Vir, Parsa Madinei
Cooking meals can be difficult, causing many to resort to cookbooks and online recipes. However, relying on these traditional methods of cooking often results in missing ingredient…