activity
20152023
most citedHow to Escape Saddle Points Efficiently

229 citations · 669 across the 20 of their papers we have counts for

collaborators
Showing 2023Show all

7 papers · 1 filter

eess.SY2023

PEBO-SLAM: Observer design for visual inertial SLAM with convergence guarantees

Bowen Yi, Chi Jin, Lei Wang +3

This paper introduces a new linear parameterization to the problem of visual inertial simultaneous localization and mapping (VI-SLAM) -- without any approximation -- for the case o…

cs.LG2023

Is RLHF More Difficult than Standard RL?

Yuanhao Wang, Qinghua Liu, Chi Jin

Reinforcement learning from Human Feedback (RLHF) learns from preference signals, while standard Reinforcement Learning (RL) directly learns from reward signals. Preferences arguab…

cs.LG2023

Context-lumpable stochastic bandits

Chung-Wei Lee, Qinghua Liu, Yasin Abbasi-Yadkori +3

We consider a contextual bandit problem with contexts and actions. In each round , the learner observes a random context and chooses an action based on its pas…

cs.LG2023

DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method

Ahmed Khaled, Konstantin Mishchenko, Chi Jin

This paper proposes a new easy-to-implement parameter-free gradient-based optimizer: DoWG (Distance over Weighted Gradients). We prove that DoWG is efficient -- matching the conver…

cs.RO2023

Learning a Universal Human Prior for Dexterous Manipulation from Human Preference

Zihan Ding, Yuanpei Chen, Allen Z. Ren +4

Generating human-like behavior on robots is a great challenge especially in dexterous manipulation tasks with robotic hands. Scripting policies from scratch is intractable due to t…

stat.ML20233 cited

On the Provable Advantage of Unsupervised Pretraining

Jiawei Ge, Shange Tang, Jianqing Fan +1

Unsupervised pretraining, which learns a useful representation using a large amount of unlabeled data to facilitate the learning of downstream tasks, is a critical component of mod…