9 citations · 21 across the 13 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
LVCHAT: Facilitating Long Video Comprehension
Yu Wang, Zeyuan Zhang, Julian McAuley +1
Enabling large language models (LLMs) to read videos is vital for multimodal LLMs. Existing works show promise on short videos whereas long video (longer than e.g.~1 minute) compre…
cs.CV2023★ 9 cited
Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models
An Yan, Yu Wang, Yiwu Zhong +8
Medical image classification is a critical problem for healthcare, with the potential to alleviate the workload of doctors and facilitate diagnoses of patients. However, two challe…
cs.CV2023★ 1 cited
Learning Concise and Descriptive Attributes for Visual Recognition
An Yan, Yu Wang, Yiwu Zhong +6
Recent advances in foundation models present new opportunities for interpretable visual recognition -- one can first query Large Language Models (LLMs) to obtain a set of attribute…