48 citations · 481 across the 104 of their papers we have counts for
19 papers · 2 filters
Deep Generative Symbolic Regression
Samuel Holt, Zhaozhi Qian, Mihaela van der Schaar
Symbolic regression (SR) aims to discover concise closed-form mathematical equations from data, a task fundamental to scientific discovery. However, the problem is highly challengi…
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
Nabeel Seedat, Nicolas Huynh, Boris van Breugel +1
Machine Learning (ML) in low-data settings remains an underappreciated yet crucial problem. Hence, data augmentation methods to increase the sample size of datasets needed for ML a…
TRIAGE: Characterizing and auditing training data for improved regression
Nabeel Seedat, Jonathan Crabbé, Zhaozhi Qian +1
Data quality is crucial for robust machine learning algorithms, with the recent interest in data-centric AI emphasizing the importance of training data characterization. However, c…
Clairvoyance: A Pipeline Toolkit for Medical Time Series
Daniel Jarrett, Jinsung Yoon, Ioana Bica +3
Time-series learning is the bread and butter of data-driven *clinical decision support*, and the recent explosion in ML research has demonstrated great potential in various healthc…
Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of Examples
Hao Sun, Alihan Hüyük, Daniel Jarrett +1
Learning controllers with offline data in decision-making systems is an essential area of research due to its potential to reduce the risk of applications in real-world systems. Ho…
Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test Data
Boris van Breugel, Nabeel Seedat, Fergus Imrie +1
Evaluating the performance of machine learning models on diverse and underrepresented subgroups is essential for ensuring fairness and reliability in real-world applications. Howev…