activity
20222026
most citedCPL: Counterfactual Prompt Learning for Vision and Language Models

3 citations · 4 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models

Yunkai Zhang, Linda Li, Yingxin Cui +5

Vision-Language Models (VLMs) excel on many multimodal reasoning benchmarks, but these evaluations often do not require an exhaustive readout of the image and can therefore obscure…

cs.CV2025

GenIR: Generative Visual Feedback for Mental Image Retrieval

Diji Yang, Minghao Liu, Chung-Hsiang Lo +2

Vision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In…

cs.CV20241 cited

Right this way: Can VLMs Guide Us to See More to Answer Questions?

Li Liu, Diji Yang, Sijia Zhong +4

In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answ…

cs.CV2024

Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution

Timothy Wei, Hsien Xin Peng, Elaine Xu +3

As Artificial Intelligence models, such as Large Video-Language models (VLMs), grow in size, their deployment in real-world applications becomes increasingly challenging due to har…

cs.CV20223 cited

CPL: Counterfactual Prompt Learning for Vision and Language Models

Xuehai He, Diji Yang, Weixi Feng +7

Prompt tuning is a new few-shot transfer learning technique that only tunes the learnable prompt for pre-trained vision and language models such as CLIP. However, existing prompt t…