3 papers
cs.CL2022
Reprint: a randomized extrapolation based on principal components for data augmentation
Le Li, Jiale Wei, Pai Peng +3
Data scarcity and data imbalance have attracted a lot of attention in many fields. Data augmentation, explored as an effective approach to tackle them, can improve the robustness a…
stat.ML2018
Sequential Learning of Principal Curves: Summarizing Data Streams on the Fly
Benjamin Guedj, Le Li
When confronted with massive data streams, summarizing data with dimension reduction methods such as PCA raises theoretical and algorithmic pitfalls. Principal curves act as a nonl…
stat.ML2016
A Quasi-Bayesian Perspective to Online Clustering
Le Li, Benjamin Guedj, Sébastien Loustau
When faced with high frequency streams of data, clustering raises theoretical and algorithmic pitfalls. We introduce a new and adaptive online clustering algorithm relying on a qua…