Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Bounded-Compute Multimodal Regression for Product-Rating Prediction
William Leach, Ru He, Sizhuo Ma +4
Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generation and dynamic visual process…
cs.CV2023
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
Ning Liao, Shaofeng Zhang, Renqiu Xia +3
There is an emerging line of research on multimodal instruction tuning, and a line of benchmarks has been proposed for evaluating these models recently. Instead of evaluating the m…