38 citations · 39 across the 7 of their papers we have counts for
11 papers · 1 filter
Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
Silviu Pitis
The softmax policy is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, explora…
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
Christopher Chiu, Silviu Pitis, Mihaela van der Schaar
Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnos…
Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
Blair Yang, Fuyang Cui, Keiran Paster +4
The rapid development and dynamic nature of large language models (LLMs) make it difficult for conventional quantitative benchmarks to accurately assess their capabilities. We prop…
Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian Rewards
Silviu Pitis
As the capabilities of artificial agents improve, they are being increasingly deployed to service multiple diverse objectives and stakeholders. However, the composition of these ob…
MoCoDA: Model-based Counterfactual Data Augmentation
Silviu Pitis, Elliot Creager, Ajay Mandlekar +1
The number of states in a dynamic process is exponential in the number of objects, making reinforcement learning (RL) difficult in complex, multi-object domains. For agents to scal…
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
Silviu Pitis, Harris Chan, Stephen Zhao +2
What goals should a multi-goal reinforcement learning agent pursue during training in long-horizon tasks? When the desired (test time) goal distribution is too distant to offer a u…