collaborators

6 papers

cs.LG2026

CleanSurvival: Automated data preprocessing for time-to-event models using reinforcement learning

Yousef Koka, David Selby, Gerrit Großmann +2

Data preprocessing is often paid little attention in machine learning, despite its potentially significant impact on model performance. While automated machine learning pipelines a…

q-bio.QM2026

A Novel Multi-view Mixture Model Framework for Longitudinal Clustering with Application to ANCA-Associated Vasculitis

Shen Jia, David Selby, Mark A Little +1

Effectively modeling irregularly sampled longitudinal data is essential for understanding disease progression and improving risk prediction. We propose a two-view mixture model tha…

cs.LG2025

X Hacking: The Threat of Misguided AutoML

Rahul Sharma, Sergey Redyuk, Sumantrak Mukherjee +4

Explainable AI (XAI) and interpretable machine learning methods help to build trust in model predictions and derived insights, yet also present a perverse incentive for analysts to…

cs.CE2025

Rethinking Cancer Gene Identification through Graph Anomaly Analysis

Yilong Zang, Lingfei Ren, Yue Li +7

Graph neural networks (GNNs) have shown promise in integrating protein-protein interaction (PPI) networks for identifying cancer genes in recent studies. However, due to the insuff…

cs.LG2025

Neural Spatiotemporal Point Processes: Trends and Challenges

Sumantrak Mukherjee, Mouad Elhamdi, George Mohler +4

Spatiotemporal point processes (STPPs) are probabilistic models for events occurring in continuous space and time. Real-world event data often exhibit intricate dependencies and he…

cs.IR2025

Had enough of experts? Quantitative knowledge retrieval from large language models

David Selby, Kai Spriestersbach, Yuichiro Iwashita +6

Large language models (LLMs) have been extensively studied for their abilities to generate convincing natural language sequences, however their utility for quantitative information…