3 papers
cs.LG2026
Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
Silviu Pitis
The softmax policy is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, explo…
cs.LG2025
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
Christopher Chiu, Silviu Pitis, Mihaela van der Schaar
Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnos…
cs.CL2024
Improving Context-Aware Preference Modeling for Language Models
Silviu Pitis, Ziang Xiao, Nicolas Le Roux +1
While finetuning language models from pairwise preferences has proven remarkably effective, the underspecified nature of natural language presents critical challenges. Direct prefe…