activity
20182022
most citedBeBold: Exploration Beyond the Boundary of Explored Regions

18 citations · 54 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CV202213 cited

Multitask Vision-Language Prompt Tuning

Sheng Shen, Shijia Yang, Tianjun Zhang +4

Prompt Tuning, conditioning on task-specific learned prompt vectors, has emerged as a data-efficient and parameter-efficient method for adapting large pretrained vision-language mo…

cs.CL20228 cited

TEMPERA: Test-Time Prompting via Reinforcement Learning

Tianjun Zhang, Xuezhi Wang, Denny Zhou +2

Careful prompt design is critical to the use of large language models in zero-shot or few-shot learning. As a consequence, there is a growing interest in automated methods to desig…

cs.LG2022

Graph Backup: Data Efficient Backup Exploiting Markovian Transitions

Zhengyao Jiang, Tianjun Zhang, Robert Kirk +2

The successes of deep Reinforcement Learning (RL) are limited to settings where we have a large stream of online experiences, but applying RL in the data-efficient setting with lim…

cs.LG20214 cited

C-Planning: An Automatic Curriculum for Learning Goal-Reaching Tasks

Tianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov +2

Goal-conditioned reinforcement learning (RL) can solve tasks in a wide range of domains, including navigation and manipulation, but learning to reach distant goals remains a centra…

cs.LG202111 cited

MADE: Exploration via Maximizing Deviation from Explored Regions

Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao +3

In online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments,…

cs.LG202018 cited

BeBold: Exploration Beyond the Boundary of Explored Regions

Tianjun Zhang, Huazhe Xu, Xiaolong Wang +4

Efficient exploration under sparse rewards remains a key challenge in deep reinforcement learning. To guide exploration, previous work makes extensive use of intrinsic reward (IR).…