2 papers
cs.LG2026
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
Haoran Zhang, Feng Zhou
Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based l…
cs.LG2026
Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding
Haoran Zhang, Chuanpu Li, Yuxin Fu +4
Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and de…