2 papers
cs.CV2026
Interpretable Modeling of Driver Attention Shifts with a Vision-Language Model
Kaiser Hamid, Khandakar Ashrafi Akbar, Peihang Li +1
Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which road object or region is being moni…
cs.CV2025
Audio-centric Video Understanding Benchmark without Text Shortcut
Yudong Yang, Jimin Zhuang, Guangzhi Sun +7
Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information.…