3 papers
cs.LG2026
Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts
Klaus-Rudolf Kladny, Maximilian Mordig, Bernhard Schölkopf +1
Mixture-of-experts (MoE) models enable scalable transformer architectures by activating only a subset of experts per token. Recent evidence suggests that performance improves with…
cs.CL2026
Rethinking Easy-to-Hard: Limits of Curriculum Learning in Post-Training for Deductive Reasoning
Maximilian Mordig, Andreas Opedal, Weiyang Liu +1
Curriculum learning (CL), motivated by the intuition that learning in increasing order of difficulty should ease generalization, is commonly adopted both in pre-training and post-t…
cs.AI2025
Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
Frederike Lübeck, Jonas Wildberger, Frederik Träuble +4
Cardiovascular disease (CVD) risk prediction models are essential for identifying high-risk individuals and guiding preventive actions. However, existing models struggle with the c…