2 papers
cs.LG2026
Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers
Kaushik Pendiyala, Haris Zia, Trevin Lee +8
Mixture-of-Experts (MoE) models can increase parameter capacity without proportionally increasing active computation, but it is unclear how this trade-off behaves in particle-physi…
hep-ph2025
Why Is Attention Sparse In Particle Transformer?
Timothy Legge, Aaron Wang, Jacob Ortiz +7
Transformer-based models have achieved state-of-the-art performance in jet tagging at the CERN Large Hadron Collider (LHC), with the Particle Transformer (ParT) representing a lead…