1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Weiqing Luo, Zhen Tan, Yifan Li +4
Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fi…