17 papers
GMGaze: MoE-Based Context-Aware Gaze Estimation with CLIP and Multiscale Transformer
Xinyuan Zhao, Yihang Wu, Ahmad Chaddad +2
Gaze estimation methods commonly use facial appearances to predict the direction of a person gaze. However, previous studies show three major challenges with convolutional neural n…
Impact of domain adaptation in deep learning for medical image classifications
Yihang Wu, Ahmad Chaddad
Domain adaptation (DA) is a quickly expanding area in machine learning that involves adjusting a model trained in one domain to perform well in another domain. While there have bee…
Deep Modeling and Interpretation for Bladder Cancer Classification
Ahmad Chaddad, Yihang Wu, Xianrui Chen
Deep models based on vision transformer (ViT) and convolutional neural network (CNN) have demonstrated remarkable performance on natural datasets. However, these models may not be…
Federated Vision Transformer with Adaptive Focal Loss for Medical Image Classification
Xinyuan Zhao, Yihang Wu, Ahmad Chaddad +2
While deep learning models like Vision Transformer (ViT) have achieved significant advances, they typically require large datasets. With data privacy regulations, access to many or…
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
Xinyuan Zhao, Xianrui Chen, Ahmad Chaddad
We present a semantics modulated, multi scale Transformer for 3D gaze estimation. Our model conditions CLIP global features with learnable prototype banks (illumination, head pose,…
Federated CLIP for Resource-Efficient Heterogeneous Medical Image Classification
Yihang Wu, Ahmad Chaddad
Despite the remarkable performance of deep models in medical imaging, they still require source data for training, which limits their potential in light of privacy concerns. Federa…