4 papers
ITBoost: Information-Theoretic Trust for Robust Boosting
Ye Su, Longlong Zhao, Diego Garcia-Gil +4
Gradient boosting remains a strong and widely used method for tabular data learning, but its performance often degrades when training labels are noisy. This behavior is largely rel…
Smart Data driven Decision Trees Ensemble Methodology for Imbalanced Big Data
Diego García-Gil, Salvador García, Ning Xiong +1
Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to trad…
DPASF: A Flink Library for Streaming Data preprocessing
Alejandro Alcalde-Barros, Diego García-Gil, Salvador García +1
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. A…
Enabling Smart Data: Noise filtering in Big Data classification
Diego García-Gil, Julián Luengo, Salvador García +1
In any knowledge discovery process the value of extracted knowledge is directly related to the quality of the data used. Big Data problems, generated by massive growth in the scale…