7 citations · 8 across the 3 of their papers we have counts for
1 paper · 1 filter
Yecheng Zhang, Rong Zhao, Zhizhou Sha +10
Vision-language models (VLMs) can describe urban scenes in rich detail, yet consistently fail to produce reliable human preference labels in domain-specific tasks such as safety as…