6 papers
CleanSurvival: Automated data preprocessing for time-to-event models using reinforcement learning
Yousef Koka, David Selby, Gerrit GroÃmann +2
Data preprocessing is often paid little attention in machine learning, despite its potentially significant impact on model performance. While automated machine learning pipelines a…
A Novel Multi-view Mixture Model Framework for Longitudinal Clustering with Application to ANCA-Associated Vasculitis
Shen Jia, David Selby, Mark A Little +1
Effectively modeling irregularly sampled longitudinal data is essential for understanding disease progression and improving risk prediction. We propose a two-view mixture model tha…
X Hacking: The Threat of Misguided AutoML
Rahul Sharma, Sergey Redyuk, Sumantrak Mukherjee +4
Explainable AI (XAI) and interpretable machine learning methods help to build trust in model predictions and derived insights, yet also present a perverse incentive for analysts to…
Rethinking Cancer Gene Identification through Graph Anomaly Analysis
Yilong Zang, Lingfei Ren, Yue Li +7
Graph neural networks (GNNs) have shown promise in integrating protein-protein interaction (PPI) networks for identifying cancer genes in recent studies. However, due to the insuff…
Neural Spatiotemporal Point Processes: Trends and Challenges
Sumantrak Mukherjee, Mouad Elhamdi, George Mohler +4
Spatiotemporal point processes (STPPs) are probabilistic models for events occurring in continuous space and time. Real-world event data often exhibit intricate dependencies and he…
Had enough of experts? Quantitative knowledge retrieval from large language models
David Selby, Kai Spriestersbach, Yuichiro Iwashita +6
Large language models (LLMs) have been extensively studied for their abilities to generate convincing natural language sequences, however their utility for quantitative information…