9 citations · 18 across the 21 of their papers we have counts for
6 papers · 1 filter
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
Fatih Ilhan, Gaowen Liu, Ramana Rao Kompella +5
Large Vision-Language Models (VLMs) have achieved remarkable success in multi-modal reasoning, but their inference time efficiency remains a significant challenge due to the memory…
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
Yichang Xu, Gaowen Liu, Ramana Rao Kompella +6
This paper presents a multi-agent perception-action exploration alliance, dubbed A4VL, for efficient long-video reasoning. A4VL operates in a multi-round perception-action explorat…
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
Selim Furkan Tekin, Yichang Xu, Gaowen Liu +3
With the growing number and diversity of Vision-Language Models (VLMs), many works explore language-based ensemble, collaboration, and routing techniques across multiple VLMs to im…
Adversarial Attention Perturbations for Large Object Detection Transformers
Zachary Yahn, Selim Furkan Tekin, Fatih Ilhan +5
Adversarial perturbations are useful tools for exposing vulnerabilities in neural networks. Existing adversarial perturbation methods for object detection are either limited to att…
A Neurosymbolic Agent System for Compositional Visual Reasoning
Yichang Xu, Gaowen Liu, Ramana Rao Kompella +5
The advancement in large language models (LLMs) and large vision models has fueled the rapid progress in multi-modal vision-language reasoning capabilities. However, existing visio…
Personalized Privacy Protection Mask Against Unauthorized Facial Recognition
Ka-Ho Chow, Sihao Hu, Tiansheng Huang +1
Face recognition (FR) can be abused for privacy intrusion. Governments, private companies, or even individual attackers can collect facial images by web scraping to build an FR sys…