15 citations · 37 across the 6 of their papers we have counts for
4 papers · 1 filter
From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models
Jiaxian Guo, Junnan Li, Dongxu Li +4
Large language models (LLMs) have demonstrated excellent zero-shot generalization to new language tasks. However, effective utilization of LLMs for zero-shot visual question-answer…
Mitigating and Evaluating Static Bias of Action Representations in the Background and the Foreground
Haoxin Li, Yuan Liu, Hanwang Zhang +1
In video action recognition, shortcut static features can interfere with the learning of motion features, resulting in poor out-of-distribution (OOD) generalization. The video back…
A Survey of Computer Vision Technologies In Urban and Controlled-environment Agriculture
Jiayun Luo, Boyang Li, Cyril Leung
In the evolution of agriculture to its next stage, Agriculture 5.0, artificial intelligence will play a central role. Controlled-environment agriculture, or CEA, is a special form…
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training
Anthony Meng Huat Tiong, Junnan Li, Boyang Li +2
Visual question answering (VQA) is a hallmark of vision and language reasoning and a challenging task under the zero-shot setting. We propose Plug-and-Play VQA (PNP-VQA), a modular…