2 papers
cs.CV2026
Distilled Large Language Model-Driven Dynamic Sparse Expert Activation Mechanism
Qinghui Chen, Zekai Zhang, Zaigui Zhang +5
High inter-class similarity, extreme scale variation, and limited computational budgets hinder reliable visual recognition across diverse real-world data. Existing vision-centric a…
cs.CV2024
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
Yuehao Song, Xinggang Wang, Jingfeng Yao +3
Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-mod…