8 papers
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces
Nicolás Astorga, Nabeel Seedat, Mihaela van der Schaar
Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making. Many real-world tasks require findin…
Skill Neologisms: Towards Skill-based Continual Learning
Antonin Berthon, Nicolas Astorga, Mihaela van der Schaar
Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable ma…
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
Claudio Fanconi, Nicolás Astorga, Mihaela van der Schaar
Teaching large language models (LLMs) to reason during post-training typically relies on reinforcement learning with explicit outcome- or process-based reward functions. However, i…
CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators
Nicolás Astorga, Anita Kriz, Mihaela van der Schaar
Despite surpassing human performance across mathematics, coding, and other knowledge-intensive tasks, large language models (LLMs) continue to struggle with causal reasoning. A cor…
Timely Clinical Diagnosis through Active Test Selection
Silas Ruhrberg Estévez, Nicolás Astorga, Mihaela van der Schaar
There is growing interest in using machine learning (ML) to support clinical diagnosis, but most approaches rely on static, fully observed datasets and fail to reflect the sequenti…
Continuously Updating Digital Twins using Large Language Models
Harry Amad, Nicolás Astorga, Mihaela van der Schaar
Digital twins are models of real-world systems that can simulate their dynamics in response to potential actions. In complex settings, the state and action variables, and available…