4 papers
Increasing Entropy to Boost Policy Gradient Performance on Personalization Tasks
Andrew Starnes, Anton Dereventsov, Clayton Webster
In this effort, we consider the impact of regularization on the diversity of actions taken by policies generated from reinforcement learning agents trained using a policy gradient.…
Zero-Shot Recommendations with Pre-Trained Large Language Models for Multimodal Nudging
Rachel M. Harrison, Anton Dereventsov, Anton Bibin
We present a method for zero-shot recommendation of multimodal non-stationary content that leverages recent advancements in the field of generative AI. We propose rendering inputs…
Modeling Non-deterministic Human Behaviors in Discrete Food Choices
Andrew Starnes, Anton Dereventsov, E. Susanne Blazek +1
We establish a non-deterministic model that predicts a user's food preferences from their demographic information. Our simulator is based on NHANES dataset and domain expert knowle…
On the Unreasonable Efficiency of State Space Clustering in Personalization Tasks
Anton Dereventsov, Ranga Raju Vatsavai, Clayton Webster
In this effort we consider a reinforcement learning (RL) technique for solving personalization tasks with complex reward signals. In particular, our approach is based on state spac…