3 papers
cs.AI2026
ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
Qiao Yan, Yihan Wang, Zhenghao Xing +2
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or…
cs.CL2025
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making
Yihan Wang, Qiao Yan, Zhenghao Xing +5
Large language models (LLMs) have demonstrated strong potential in clinical question answering, with recent multi-agent frameworks further improving diagnostic accuracy via collabo…
cs.AI2025
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
Qiao Yan, Yuchen Yuan, Xiaowei Hu +6
The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the models may generate seemingly plaus…