Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
Qiao Yan, Yihan Wang, Zhenghao Xing +2
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or…
cs.AI2025
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
Qiao Yan, Yuchen Yuan, Xiaowei Hu +6
The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the models may generate seemingly plaus…