4.3k citations · 12.1k across the 42 of their papers we have counts for
6 papers · 1 filter
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes +9
While Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to prompt injection attacks: malicious third party prompts that su…
Dropout Reduces Underfitting
Zhuang Liu, Zhiqiu Xu, Joseph Jin +2
Introduced by Hinton et al. in 2012, dropout has stood the test of time as a regularizer for preventing overfitting in neural networks. In this study, we demonstrate that dropout c…
Explaining Reinforcement Learning Policies through Counterfactual Trajectories
Julius Frost, Olivia Watkins, Eric Weiner +4
In order for humans to confidently decide where to employ RL agents for real-world tasks, a human developer must validate that the agent will perform well at test-time. Some policy…
Loss is its own Reward: Self-Supervision for Reinforcement Learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus +1
Reinforcement learning optimizes policies for expected cumulative reward. Need the supervision be so narrow? Reward is delayed and sparse for many tasks, making it a difficult and…
Learning Modular Neural Network Policies for Multi-Task and Multi-Robot Transfer
Coline Devin, Abhishek Gupta, Trevor Darrell +2
Reinforcement learning (RL) can automate a wide variety of robotic skills, but learning each new skill requires considerable real-world data collection and manual representation en…
Multi-View Learning in the Presence of View Disagreement
C. Christoudias, Raquel Urtasun, Trevor Darrell
Traditional multi-view learning approaches suffer in the presence of view disagreement,i.e., when samples in each view do not belong to the same class due to view corruption, occlu…