2 papers
cs.LG2026
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output
Guozheng Li, Xiyan Fu, Yiwen Guo
Current reinforcement learning from human feedback (RLHF) methods primarily rely on scalar rewards from a trained reward model (RM). While effective, scalar rewards are often noisy…
cs.HC2024
HiRegEx: Interactive Visual Query and Exploration of Multivariate Hierarchical Data
Guozheng Li, Haotian Mi, Chi Harold Liu +2
When using exploratory visual analysis to examine multivariate hierarchical data, users often need to query data to narrow down the scope of analysis. However, formulating effectiv…