activity
20232026
most citedWhy are Visually-Grounded Language Models Bad at Image Classification?

4 citations · 7 across the 16 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

6 papers · 2 filters

cs.CV2024

Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Mark Endo, Xiaohan Wang, Serena Yeung-Levy

Recent works on accelerating Vision-Language Models achieve strong performance across a variety of vision-language tasks despite highly compressing visual information. In this work…

cs.CV2024

Zero-shot Action Localization via the Confidence of Large Vision-Language Models

Josiah Aklilu, Xiaohan Wang, Serena Yeung-Levy

Precise action localization in untrimmed video is vital for fields such as professional sports and minimally invasive surgery, where the delineation of particular motions in record…

cs.CV2024

Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Orr Zohar, Xiaohan Wang, Yonatan Bitton +2

The performance of Large Vision Language Models (LVLMs) is dependent on the size and quality of their training datasets. Existing video instruction tuning datasets lack diversity a…

cs.CV2024

Why are Visually-Grounded Language Models Bad at Image Classification?

Yuhui Zhang, Alyssa Unell, Xiaohan Wang +4

Image classification is one of the most fundamental capabilities of machine vision intelligence. In this work, we revisit the image classification task using visually-grounded lang…

cs.CV20241 cited

VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Xiaohan Wang, Yuhui Zhang, Orr Zohar +1

Long-form video understanding represents a significant challenge within computer vision, demanding a model capable of reasoning over long multi-modal sequences. Motivated by the hu…

cs.CV2024

Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models

Elaine Sui, Xiaohan Wang, Serena Yeung-Levy

Advancements in vision-language models (VLMs) have propelled the field of computer vision, particularly in the zero-shot learning setting. Despite their promise, the effectiveness…