4 papers
GMGaze: MoE-Based Context-Aware Gaze Estimation with CLIP and Multiscale Transformer
Xinyuan Zhao, Yihang Wu, Ahmad Chaddad +2
Gaze estimation methods commonly use facial appearances to predict the direction of a person gaze. However, previous studies show three major challenges with convolutional neural n…
Federated Vision Transformer with Adaptive Focal Loss for Medical Image Classification
Xinyuan Zhao, Yihang Wu, Ahmad Chaddad +2
While deep learning models like Vision Transformer (ViT) have achieved significant advances, they typically require large datasets. With data privacy regulations, access to many or…
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
Xinyuan Zhao, Xianrui Chen, Ahmad Chaddad
We present a semantics modulated, multi scale Transformer for 3D gaze estimation. Our model conditions CLIP global features with learnable prototype banks (illumination, head pose,…
A Knowledge Distillation-Based Approach to Enhance Transparency of Classifier Models
Yuchen Jiang, Xinyuan Zhao, Yihang Wu +1
With the rapid development of artificial intelligence (AI), especially in the medical field, the need for its explainability has grown. In medical image analysis, a high degree of…