7 papers
Cluster Analysis with Resampling for Validation and Exploration (CARVE)
Kai R. Wycik, Tiffany M. Tang, Tarek M. Zikry +1
Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries. However, clustering results are highly sensitive to the choice of…
Estimating Consensus Ideal Points Using Multi-Source Data
Mellissa Meisels, Melody Huang, Tiffany M. Tang
In the advent of big data and machine learning, researchers now have a wealth of congressional candidate ideal point estimates at their disposal for theory testing. Weak relationsh…
Consensus dimension reduction via multi-view learning
Bingxue An, Tiffany M. Tang
A plethora of dimension reduction methods have been developed to visualize high-dimensional data in low dimensions. However, different dimension reduction methods often output diff…
Interpretable Network-assisted Random Forest+
Tiffany M. Tang, Elizaveta Levina, Ji Zhu
Machine learning algorithms often assume that training samples are independent. When data points are connected by a network, the induced dependency between samples is both a challe…
Top- Feature Importance Ranking
Yuxi Chen, Tiffany Tang, Genevera Allen
Accurate ranking of important features is a fundamental challenge in interpretable machine learning with critical applications in scientific discovery and decision-making. Unlike f…
Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices
Andersen Chang, Tiffany M. Tang, Tarek M. Zikry +1
Unsupervised machine learning is widely used to mine large, unlabeled datasets to make data-driven discoveries in critical domains such as climate science, biomedicine, astronomy,…