3 papers
cs.CV2026
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
Athul M. Mathew, Haithem Hermassi, Thariq Khalid +1
Gaze understanding unifies the detection of people, their gaze targets, and objects of interest into a single framework, offering critical insight into visual attention and intent…
cs.CL2025
Increasing the Thinking Budget is Not All You Need
Ignacio Iacobacci, Zhaozhi Qian, Faroq AL-Tam +2
Recently, a new wave of thinking-capable Large Language Models has emerged, demonstrating exceptional capabilities across a wide range of reasoning benchmarks. Early studies have b…
cs.CV2025
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
Athul M. Mathew, Arshad Ali Khan, Thariq Khalid +2
Gaze target detection (GTD) is the task of predicting where a person in an image is looking. This is a challenging task, as it requires the ability to understand the relationship b…