Mixture-of-Experts Graph Transformers for Interpretable Particle Collision Detection
arXiv:2501.03432 · doi:10.1038/s41598-025-12003-9
Abstract
The Large Hadron Collider at CERN produces immense volumes of complex data from high-energy particle collisions, demanding sophisticated analytical techniques for effective interpretation. Neural Networks, including Graph Neural Networks, have shown promise in tasks such as event classification and object identification by representing collisions as graphs. However, while Graph Neural Networks excel in predictive accuracy, their "black box" nature often limits their interpretability, making it difficult to trust their decision-making processes. In this paper, we propose a novel approach that combines a Graph Transformer model with Mixture-of-Expert layers to achieve high predictive performance while embedding interpretability into the architecture. By leveraging attention maps and expert specialization, the model offers insights into its internal decision-making, linking predictions to physics-informed features. We evaluate the model on simulated events from the ATLAS experiment, focusing on distinguishing rare Supersymmetric signal events from Standard Model background. Our results highlight that the model achieves competitive classification accuracy while providing interpretable outputs that align with known physics, demonstrating its potential as a robust and transparent tool for high-energy physics data analysis. This approach underscores the importance of explainability in machine learning methods applied to high energy physics, offering a path toward greater trust in AI-driven discoveries.
References in corpus (20)
- XGBoost: A Scalable Tree Boosting System
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- GNNExplainer: Generating Explanations for Graph Neural Networks
- Graph Learning: A Survey
- Deep Learning and its Application to LHC Physics
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- A Generalization of Transformer Networks to Graphs
- Parameterized Explainer for Graph Neural Network
- Graph Neural Networks in Particle Physics
- Estimating Training Data Influence by Tracing Gradient Descent
- Graph Neural Networks for Particle Reconstruction in High Energy Physics detectors
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Graph Neural Networks for Particle Tracking and Reconstruction
- Sign and Basis Invariant Networks for Spectral Graph Representation Learning
- Software Performance of the ATLAS Track Reconstruction for LHC Run 3
- Graph Neural Networks in Particle Physics: Implementations, Innovations, and Challenges
- Do Feature Attribution Methods Correctly Attribute Features?
- Transport Simulation and Diffractive Event Reconstruction at the LHC
- Search for direct production of electroweakinos in final states with one lepton, jets and missing transverse momentum in pp collisions at TeV with the ATLAS detector