2 citations · 8 across the 42 of their papers we have counts for
3 papers · 1 filter
Many Preferences, Few Policies: Towards Scalable Language Model Personalization
Cheol Woo Kim, Jai Moondra, Roozbeh Nahavandi +3
The holy grail of LLM personalization is a single LLM for each user, perfectly aligned with that user's preferences. However, maintaining a separate LLM per user is impractical due…
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
Xinyi Yang, Liang Zeng, Heng Dong +6
As humans increasingly share environments with diverse agents powered by RL, LLMs, and beyond, the ability to explain agent policies in natural language is vital for reliable coexi…
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
Ambreesh Parthasarathy, Chandrasekar Subramanian, Ganesh Senrayan +4
Restless Multi-Armed Bandits (RMABs) have been successfully applied to resource allocation problems in a variety of settings, including public health. With the rapid development of…