10 citations · 11 across the 3 of their papers we have counts for
4 papers
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
Tony Lee, Andrew Wagenmaker, Karl Pertsch +3
A well-designed reward is critical for effective reinforcement learning-based policy improvement. In real-world robotics, obtaining such rewards typically requires either labor-int…
AHELM: A Holistic Evaluation of Audio-Language Models
Tony Lee, Haoqin Tu, Chi Heem Wong +6
Evaluations of audio-language models (ALMs) -- multimodal models that take interleaved audio and text as input and output text -- are hindered by the lack of standardized benchmark…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
Pranav Atreya, Karl Pertsch, Tony Lee +29
Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standar…