3 citations · 3 across the 3 of their papers we have counts for
3 papers
When Can Proxies Improve the Sample Complexity of Preference Learning?
Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi +4
We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as…
Plain-Det: A Plain Multi-Dataset Object Detector
Cheng Shi, Yuchen Zhu, Sibei Yang
Recent advancements in large-scale foundational models have sparked widespread interest in training highly proficient large vision models. A common consensus revolves around the ne…
Deep Learning for Mean Field Optimal Transport
Sebastian Baudelet, Brieuc Frénais, Mathieu Laurière +2
Mean field control (MFC) problems have been introduced to study social optima in very large populations of strategic agents. The main idea is to consider an infinite population and…