activity
20162025
most citedDesigning Rewards for Fast Learning

8 citations · 13 across the 8 of their papers we have counts for

collaborators

12 papers

cs.CL2025

Generalist Reward Models: Found Inside Large Language Models

Yi-Chen Li, Tian Xu, Yang Yu +6

The alignment of Large Language Models (LLMs) is critically dependent on reward models trained on costly human preference data. While recent work explores bypassing this cost with…

cs.AR2025★ 1 cited

Open3DBench: Open-Source Benchmark for 3D-IC Backend Implementation and PPA Evaluation

Yunqi Shi, Chengrui Gao, Wanqi Ren +6

This work introduces Open3DBench, an open-source 3D-IC backend implementation benchmark built upon the OpenROAD-flow-scripts framework, enabling comprehensive evaluation of power,…

cs.LG2024★ 1 cited

WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making

Zhilong Zhang, Ruifeng Chen, Junyin Ye +8

World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate…

cs.SE2024

Beimingwu: A Learnware Dock System

Zhi-Hao Tan, Jian-Dong Liu, Xiao-Dong Bi +7

The learnware paradigm proposed by Zhou [2016] aims to enable users to reuse numerous existing well-trained models instead of building machine learning models from scratch, with th…

cs.MA2023

Learning to Coordinate with Anyone

Lei Yuan, Lihe Li, Ziqian Zhang +5

In open multi-agent environments, the agents may encounter unexpected teammates. Classical multi-agent learning approaches train agents that can only coordinate with seen teammates…

cs.LG2022★ 8 cited

Designing Rewards for Fast Learning

Henry Sowerby, Zhiyuan Zhou, Michael L. Littman

To convey desired behavior to a Reinforcement Learning (RL) agent, a designer must choose a reward function for the environment, arguably the most important knob designers have in…