1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Trong Thang Pham, Hien Nguyen, Ngan Le
Current multimodal large language models (MLLMs) cannot effectively utilize eye-gaze information for video understanding, even when gaze cues are supplied via visual overlays or te…