8 citations · 13 across the 8 of their papers we have counts for
12 papers
Generalist Reward Models: Found Inside Large Language Models
Yi-Chen Li, Tian Xu, Yang Yu +6
The alignment of Large Language Models (LLMs) is critically dependent on reward models trained on costly human preference data. While recent work explores bypassing this cost with…
Open3DBench: Open-Source Benchmark for 3D-IC Backend Implementation and PPA Evaluation
Yunqi Shi, Chengrui Gao, Wanqi Ren +6
This work introduces Open3DBench, an open-source 3D-IC backend implementation benchmark built upon the OpenROAD-flow-scripts framework, enabling comprehensive evaluation of power,…
WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making
Zhilong Zhang, Ruifeng Chen, Junyin Ye +8
World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate…
Beimingwu: A Learnware Dock System
Zhi-Hao Tan, Jian-Dong Liu, Xiao-Dong Bi +7
The learnware paradigm proposed by Zhou [2016] aims to enable users to reuse numerous existing well-trained models instead of building machine learning models from scratch, with th…
Learning to Coordinate with Anyone
Lei Yuan, Lihe Li, Ziqian Zhang +5
In open multi-agent environments, the agents may encounter unexpected teammates. Classical multi-agent learning approaches train agents that can only coordinate with seen teammates…
Designing Rewards for Fast Learning
Henry Sowerby, Zhiyuan Zhou, Michael L. Littman
To convey desired behavior to a Reinforcement Learning (RL) agent, a designer must choose a reward function for the environment, arguably the most important knob designers have in…