7 papers
Multi-Robot Coordination for Planning under Context Uncertainty
Pulkit Rustagi, Kyle Hollins Wray, Sandhya Saisubramanian
Real-world robots often operate in settings where objective priorities depend on the underlying context of operation. When the underlying context is unknown apriori, multiple robot…
Active teacher selection for reward learning
Rachel Freedman, Justin Svegliato, Kyle Wray +1
Reward learning techniques enable machine learning systems to learn objectives from human feedback. A core limitation of these systems is their assumption that all feedback comes f…
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
Saaduddin Mahmud, Mason Nakamura, Kyle Hollins Wray +1
Prompt optimization methods have demonstrated significant effectiveness in aligning black-box large language models (LLMs). In parallel, inference scaling strategies such as Best-o…
Multi-Objective Multi-Agent Path Finding with Lexicographic Cost Preferences
Pulkit Rustagi, Kyle Hollins Wray, Sandhya Saisubramanian
Many real-world scenarios require multiple agents to coordinate in shared environments, while balancing trade-offs between multiple, potentially competing objectives. Current multi…
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
Mason Nakamura, Saaduddin Mahmud, Kyle H. Wray +2
Aligning LLMs with user preferences is crucial for real-world use but often requires costly fine-tuning or expensive inference, forcing trade-offs between alignment quality and com…
Rao-Blackwellized POMDP Planning
Jiho Lee, Nisar R. Ahmed, Kyle H. Wray +1
Partially Observable Markov Decision Processes (POMDPs) provide a structured framework for decision-making under uncertainty, but their application requires efficient belief update…