Publications (9)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
Anas Barakat, Souradip Chakraborty, Peihong Yu +2
Reinforcement learning with general utilities (RLGU) offers a unifying framework to capture several problems beyond standard expected returns, including imitation learning, pure ex…
TACTIC: Task-Agnostic Contrastive pre-Training for Inter-Agent Communication
Peihong Yu, Manav Mishra, Syed Zaidi +1
The "sight range dilemma" in cooperative Multi-Agent Reinforcement Learning (MARL) presents a significant challenge: limited observability hinders team coordination, while extensiv…
Enhancing Multi-Agent Coordination through Common Operating Picture Integration
Peihong Yu, Bhoram Lee, Aswin Raghavan +3
In multi-agent systems, agents possess only local observations of the environment. Communication between teammates becomes crucial for enhancing coordination. Past research has pri…
Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings
Keenan Powell, Peihong Yu, Pratap Tokekar
Many Multi-Agent Reinforcement Learning (MARL) agents fail to adapt properly to cooperating with agents trained with the same objectives but different seeds, algorithms, or other t…
Beyond Joint Demonstrations: Personalized Expert Guidance for Efficient Multi-Agent Reinforcement Learning
Peihong Yu, Manav Mishra, Alec Koppel +5
Multi-Agent Reinforcement Learning (MARL) algorithms face the challenge of efficient exploration due to the exponential increase in the size of the joint state-action space. While…
Insta-RS: Instance-wise Randomized Smoothing for Improved Robustness and Accuracy
Chen Chen, Kezhi Kong, Peihong Yu +3
Randomized smoothing (RS) is an effective and scalable technique for constructing neural network classifiers that are certifiably robust to adversarial perturbations. Most RS works…
VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic Labels
Chen Chen, Hannah K. Bako, Peihong Yu +10
Chart corpora, which comprise data visualizations and their semantic labels, are crucial for advancing visualization research. However, the labels in most existing corpora are high…
Sketch-to-Skill: Bootstrapping Robot Learning with Human Drawn Trajectory Sketches
Peihong Yu, Amisha Bhaskar, Anukriti Singh +2
Training robotic manipulation policies traditionally requires numerous demonstrations and/or environmental rollouts. While recent Imitation Learning (IL) and Reinforcement Learning…
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
Anukriti Singh, Amisha Bhaskar, Peihong Yu +4
Designing reward functions for continuous-control robotics often leads to subtle misalignments or reward hacking, especially in complex tasks. Preference-based RL mitigates some of…