activity
20092023
most citedIntegrative Generalized Convex Clustering Optimization and Feature Selection for Mixed Multi-View Data

21 citations · 25 across the 11 of their papers we have counts for

collaborators
Showing stat.MLShow all

8 papers · 1 filter

stat.ML2023

Data Augmentation via Subgroup Mixup for Improving Fairness

Madeline Navarro, Camille Little, Genevera I. Allen +1

In this work, we propose data augmentation via pairwise mixup across subgroups to improve group fairness. Many real-world applications of machine learning systems exhibit biases ac…

stat.ML2023

Interpretable Machine Learning for Discovery: Statistical Challenges \& Opportunities

Genevera I. Allen, Luqin Gan, Lili Zheng

New technologies have led to vast troves of large and complex datasets across many scientific domains and industries. People routinely use machine learning techniques to not only p…

stat.ML20213 cited

Thresholded Graphical Lasso Adjusts for Latent Variables: Application to Functional Neural Connectivity

Minjie Wang, Genevera I. Allen

In neuroscience, researchers seek to uncover the connectivity of neurons from large-scale neural recordings or imaging; often people employ graphical model selection and estimation…

stat.ML2020

Simultaneous Grouping and Denoising via Sparse Convex Wavelet Clustering

Michael Weylandt, T. Mitchell Roddenberry, Genevera I. Allen

Clustering is a ubiquitous problem in data science and signal processing. In many applications where we observe noisy signals, it is common practice to first denoise the data, perh…

stat.ML2020

MP-Boost: Minipatch Boosting via Adaptive Feature and Observation Sampling

Mohammad Taha Toghani, Genevera I. Allen

Boosting methods are among the best general-purpose and off-the-shelf machine learning approaches, gaining widespread popularity. In this paper, we seek to develop a boosting metho…

stat.ML2020

Feature Selection for Huge Data via Minipatch Learning

Tianyi Yao, Genevera I. Allen

Feature selection often leads to increased model interpretability, faster computation, and improved model performance by discarding irrelevant or redundant features. While feature…