2 papers
cs.LG2024
On Designing Effective RL Reward at Training Time for LLM Reasoning
Jiaxuan Gao, Shusheng Xu, Wenjie Ye +6
Reward models have been increasingly critical for improving the reasoning capability of LLMs. Existing research has shown that a well-trained reward model can substantially improve…
cs.CV2023
What Happened 3 Seconds Ago? Inferring the Past with Thermal Imaging
Zitian Tang, Wenjie Ye, Wei-Chiu Ma +1
Inferring past human motion from RGB images is challenging due to the inherent uncertainty of the prediction problem. Thermal images, on the other hand, encode traces of past human…