activity
20242026
collaborators

5 papers

cs.AI2026

Measuring Dataset Diversity from a Geometric Perspective

Yang Ba, Mohammad Sadeq Abolhasani, Michelle V Mancenido +1

Diversity can be broadly defined as the presence of meaningful variation across elements, which can be viewed from multiple perspectives, including statistical variation and geomet…

cs.LG2025

Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?

Yang Ba, Michelle V. Mancenido, Rong Pan

Data augmentation that introduces diversity into the input data has long been used in training deep learning models. It has demonstrated benefits in improving robustness and genera…

cs.HC2025

Data Quality in Crowdsourcing and Spamming Behavior Detection

Yang Ba, Michelle V. Mancenido, Erin K. Chiou +1

As crowdsourcing emerges as an efficient and cost-effective method for obtaining labels for machine learning datasets, it is important to assess the quality of crowd-provided data,…

cs.CY2025

PADTHAI-MM: Principles-based Approach for Designing Trustworthy, Human-centered AI using MAST Methodology

Myke C. Cohen, Nayoung Kim, Yang Ba +7

Despite an extensive body of literature on trust in technology, designing trustworthy AI systems for high-stakes decision domains remains a significant challenge, further compounde…

cs.CL2024

Fill In The Gaps: Model Calibration and Generalization with Synthetic Data

Yang Ba, Michelle V. Mancenido, Rong Pan

As machine learning models continue to swiftly advance, calibrating their performance has become a major concern prior to practical and widespread implementation. Most existing cal…