2 papers
cs.CV2026
Deep Modeling and Interpretation for Bladder Cancer Classification
Ahmad Chaddad, Yihang Wu, Xianrui Chen
Deep models based on vision transformer (ViT) and convolutional neural network (CNN) have demonstrated remarkable performance on natural datasets. However, these models may not be…
cs.CV2026
GazeFormer-MoE: Context-Aware Gaze Estimation via CLIP and MoE Transformer
Xinyuan Zhao, Xianrui Chen, Ahmad Chaddad
We present a semantics modulated, multi scale Transformer for 3D gaze estimation. Our model conditions CLIP global features with learnable prototype banks (illumination, head pose,…