Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following
Shijing Wang, Yaping Huang, Chaoqun Cui +4
Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation models (VFMs) have demonstrated s…
cs.CV2026
CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
Liangbin Huang, Xiaohua Liao, Chaoqun Cui +4
Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic condit…
cs.CV2025
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
Shijing Wang, Chaoqun Cui, Yaping Huang +2
Human gaze provides essential cues for interpreting attention, intention, and social interaction in visual scenes, yet gaze understanding remains largely unexplored in current visi…