Detecting Variability in Massive Astronomical Time-Series Data I: application of an infinite Gaussian mixture model
arXiv:0908.2664 · doi:10.1111/j.1365-2966.2009.15576.x
Abstract
We present a new framework to detect various types of variable objects within massive astronomical time-series data. Assuming that the dominant population of objects is non-variable, we find outliers from this population by using a non-parametric Bayesian clustering algorithm based on an infinite GaussianMixtureModel (GMM) and the Dirichlet Process. The algorithm extracts information from a given dataset, which is described by six variability indices. The GMM uses those variability indices to recover clusters that are described by six-dimensional multivariate Gaussian distributions, allowing our approach to consider the sampling pattern of time-series data, systematic biases, the number of data points for each light curve, and photometric quality. Using the Northern Sky Variability Survey data, we test our approach and prove that the infinite GMM is useful at detecting variable objects, while providing statistical inference estimation that suppresses false detection. The proposed approach will be effective in the exploration of future surveys such as GAIA, Pan-Starrs, and LSST, which will produce massive time-series data.
accepted for publication in MNRAS
References in corpus (5)
- Automated supervised classification of variable stars I. Methodology
- Statistical Evidence for Three classes of Gamma-ray Bursts
- Variable stars across the observational HR diagram
- Revealing components of the galaxy population through nonparametric techniques
- Automated Probabilistic Classification of Transients and Variables
Cited by in corpus (26)
- On Machine-Learned Classification of Variable Stars with Sparse and Noisy Time-Series Data
- Comparative performance of selected variability detection techniques in photometric time series
- A comparison of period finding algorithms
- QSO Selection Algorithm Using Time Variability and Machine Learning: Selection of 1,620 QSO Candidates from MACHO LMC Database
- Construction of a Calibrated Probabilistic Classification Catalog: Application to 50k Variable Sources in the All-Sky Automated Survey
- The EPOCH Project: I. Periodic variable stars in the EROS-2 LMC database
- Machine learning search for variable stars
- A Comprehensive Power Spectral Density Analysis of Astronomical Time Series. II. The Swift/BAT Long Gamma-Ray Bursts
- A Machine Learning Classifier for Microlensing in Wide-Field Surveys
- New Insights into Time Series Analysis - I - Correlated observations
- Probing Gas Motions in the Intra-Cluster Medium: A Mixture Model Approach
- Statistical Searches for Microlensing Events in Large, Non-Uniformly Sampled Time-Domain Surveys: A Test Using Palomar Transient Factory Data
- An Automated tool to detect variable sources in the Vista Variables in the V\'ıa Láctea Survey. The VVV Variables (V) catalog of tiles d001 and d002
- The Structure of the Young Star Cluster NGC 6231. I. Stellar Population
- Detecting Variability in Massive Astronomical Time-Series Data II: Variable Candidates in the Northern Sky Variability Survey
- Classification of pulsars with Dirichlet process Gaussian mixture model
- The Hubble Catalog of Variables (HCV)
- Connecting the time domain community with the Virtual Astronomical Observatory
- Analytical representation of Gaussian processes in the plane
- Recovering variable stars in large surveys: EA Algol-type class in the Catalina Survey
- Automatic Catalog of RRLyrae from 14 million VVV Light Curves: How far can we go with traditional machine-learning?
- Variability search in M 31 using Principal Component Analysis and the Hubble Source Catalog
- A new approach to feature-based asteroid taxonomy in 3D color space: 1. SDSS photometric system
- Searching for Quasi-Periodic Eruptions using Machine Learning
- The Hubble Catalog of Variables
- The incidence of LBV variability in the LMC