10 citations · 14 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 10 cited
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman +6
Sparse autoencoders provide a promising unsupervised approach for extracting interpretable features from a language model by reconstructing activations from a sparse bottleneck lay…
cs.AI2023★ 4 cited
Action-Quantized Offline Reinforcement Learning for Robotic Skill Learning
Jianlan Luo, Perry Dong, Jeffrey Wu +3
The offline reinforcement learning (RL) paradigm provides a general recipe to convert static behavior datasets into policies that can perform better than the policy that collected…