6 papers
Nearly Optimal Bayesian Inference for Structural Missingness
Chen Liang, Donghua Yang, Yutong Zhao +9
Structural missingness breaks 'just impute and train': values can be undefined by causal or logical constraints, and the mask may depend on observed variables, unobserved variables…
Exploring the Heterogeneity of Tabular Data: A Diversity-aware Data Generator via LLMs
Yafeng Tang, Xiaoou Ding, Jianzhuo Du +5
Tabular data generation has become increasingly essential for enabling robust machine learning applications, which require large-scale, high-quality data. Existing solutions levera…
KDSelector: A Knowledge-Enhanced and Data-Efficient Model Selector Learning Framework for Time Series Anomaly Detection
Zhiyu Liang, Dongrui Cai, Chenyuan Zhang +6
Model selection has been raised as an essential problem in the area of time series anomaly detection (TSAD), because there is no single best TSAD model for the highly heterogeneous…
Revisiting Data Analysis with Pre-trained Foundation Models
Chen Liang, Donghua Yang, Zheng Liang +6
Data analysis focuses on harnessing advanced statistics, programming, and machine learning techniques to extract valuable insights from vast datasets. An increasing volume and vari…
Exploring Data and Knowledge combined Anomaly Explanation of Multivariate Industrial Data
Xiaoou Ding, Hongzhi Wang, Chen Wang +2
The demand for high-performance anomaly detection techniques of IoT data becomes urgent, especially in industry field. The anomaly identification and explanation in time series dat…
Auto-CASH: Autonomous Classification Algorithm Selection with Deep Q-Network
Tianyu Mu, Hongzhi Wang, Chunnan Wang +1
The great amount of datasets generated by various data sources have posed the challenge to machine learning algorithm selection and hyperparameter configuration. For a specific mac…