activity
20152026
most citedDeep Counterfactual Networks with Propensity-Dropout

48 citations · 236 across the 62 of their papers we have counts for

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces

Nicolás Astorga, Nabeel Seedat, Mihaela van der Schaar

Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making. Many real-world tasks require findin…

cs.AI2026

CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators

Nicolás Astorga, Anita Kriz, Mihaela van der Schaar

Despite surpassing human performance across mathematics, coding, and other knowledge-intensive tasks, large language models (LLMs) continue to struggle with causal reasoning. A cor…

cs.AI2026

TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization

Seokhyun Lee, Jaeho Kim, Changjun Oh +2

Time-series generative models often lack control over temporal granularity, forcing users to accept whatever granularity the model produces. To enable truly user-driven generation,…

cs.AI2025

Timely Clinical Diagnosis through Active Test Selection

Silas Ruhrberg Estévez, Nicolás Astorga, Mihaela van der Schaar

There is growing interest in using machine learning (ML) to support clinical diagnosis, but most approaches rely on static, fully observed datasets and fail to reflect the sequenti…

cs.AI2025

Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning

Claudio Fanconi, Nicolás Astorga, Mihaela van der Schaar

Teaching large language models (LLMs) to reason during post-training typically relies on reinforcement learning with explicit outcome- or process-based reward functions. However, i…

cs.AI2025

Preference Learning for AI Alignment: a Causal Perspective

Katarzyna Kobalczyk, Mihaela van der Schaar

Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs.…