From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Qiucheng Yu, Ruijie Xu, Mingang Chen +2
The paper introduces TSHA, a large benchmark of real-world indoor safety hazard assessment questions for evaluating vision‑language models, and shows that training on this data imp…
cs.CV2026
Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
Xianke Chen, Daizong Liu, Yushuo Lou +5
Different from traditional text-to-image retrieval tasks, chat-based image retrieval allows the human-interactive system to iteratively clarify and refine user intent through multi…
cs.LG2026
Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models
Yicheng Bao, Xuhong Wang, Qiaosheng Zhang +3
Despite their impressive capabilities, Multimodal Large Language Models (MLLMs) exhibit perceptual fragility when confronted with visually complex scenes. This weakness stems from…