Resonant Anomaly Detection with Multiple Reference Datasets
arXiv:2212.10579 · doi:10.1007/JHEP07(2023)188
Abstract
An important class of techniques for resonant anomaly detection in high energy physics builds models that can distinguish between reference and target datasets, where only the latter has appreciable signal. Such techniques, including Classification Without Labels (CWoLa) and Simulation Assisted Likelihood-free Anomaly Detection (SALAD) rely on a single reference dataset. They cannot take advantage of commonly-available multiple datasets and thus cannot fully exploit available information. In this work, we propose generalizations of CWoLa and SALAD for settings where multiple reference datasets are available, building on weak supervision techniques. We demonstrate improved performance in a number of settings with realistic and synthetic data. As an added benefit, our generalizations enable us to provide finite-sample guarantees, improving on existing asymptotic analyses.
References in corpus (8)
- Classification without labels: Learning from mixed samples in high energy physics
- Extending the Bump Hunt with Machine Learning
- CURTAINs for your Sliding Window: Constructing Unobserved Regions by Transforming Adjacent Intervals
- Quantum Anomaly Detection for Collider Physics
- Self-supervised Anomaly Detection for New Physics
- Learning new physics efficiently with nonparametric methods
- Anomaly Detection under Coordinate Transformations
- Simulation-based Anomaly Detection for Multileptons at the LHC
Cited by in corpus (4)
- The Interplay of Machine Learning--based Resonant Anomaly Detection Methods
- Anomalies, Representations, and Self-Supervision
- Combining Resonant and Tail-based Anomaly Detection
- Weakly supervised anomaly detection for resonant new physics in the dijet final state using proton-proton collisions at TeV with the ATLAS detector