48 citations · 198 across the 61 of their papers we have counts for
20 papers · 1 filter
Deep Generative Symbolic Regression
Samuel Holt, Zhaozhi Qian, Mihaela van der Schaar
Symbolic regression (SR) aims to discover concise closed-form mathematical equations from data, a task fundamental to scientific discovery. However, the problem is highly challengi…
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
Nabeel Seedat, Nicolas Huynh, Boris van Breugel +1
Machine Learning (ML) in low-data settings remains an underappreciated yet crucial problem. Hence, data augmentation methods to increase the sample size of datasets needed for ML a…
A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning
Alicia Curth, Alan Jeffares, Mihaela van der Schaar
Conventional statistical wisdom established a well-understood relationship between model complexity and prediction error, typically presented as a U-shaped curve reflecting a trans…
TRIAGE: Characterizing and auditing training data for improved regression
Nabeel Seedat, Jonathan Crabbé, Zhaozhi Qian +1
Data quality is crucial for robust machine learning algorithms, with the recent interest in data-centric AI emphasizing the importance of training data characterization. However, c…
Explaining by Imitating: Understanding Decisions by Interpretable Policy Learning
Alihan Hüyük, Daniel Jarrett, Mihaela van der Schaar
Understanding human behavior from observed data is critical for transparency and accountability in decision-making. Consider real-world settings such as healthcare, in which modeli…
Clairvoyance: A Pipeline Toolkit for Medical Time Series
Daniel Jarrett, Jinsung Yoon, Ioana Bica +3
Time-series learning is the bread and butter of data-driven *clinical decision support*, and the recent explosion in ML research has demonstrated great potential in various healthc…