10 papers
Active Real-World Factor-Based Evaluation for Generalist Robot Policies
Andrew Liao, Hanchen Cui, Karthik Desingh +1
Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks. However, rigorously evaluating these policies…
Aligning Language Models with Selective Prediction
Gaoxiang Luo, Yifan Wu, Sinian Zhang +2
Large language models (LLMs) are increasingly deployed as critical decision-making components in high-stakes real-world AI systems, rendering LLM reliability a foremost practical c…
Discovery of Feasible 3D Printing Configurations for Metal Alloys via AI-driven Adaptive Experimental Design
Azza Fadhel, Nathaniel W. Zuckschwerdt, Aryan Deshwal +3
Configuring the parameters of additive manufacturing processes for metal alloys is a challenging problem due to complex relationships between input parameters (e.g., laser power, s…
An Exploratory Study of Bayesian Prompt Optimization for Test-Driven Code Generation with Large Language Models
Shlok Tomar, Aryan Deshwal, Ethan Villalovoz +3
We consider the task of generating functionally correct code using large language models (LLMs). The correctness of generated code is influenced by the prompt used to query the giv…
Online Optimization for Offline Safe Reinforcement Learning
Yassine Chemingui, Aryan Deshwal, Alan Fern +2
We study the problem of Offline Safe Reinforcement Learning (OSRL), where the goal is to learn a reward-maximizing policy from fixed data under a cumulative cost constraint. We pro…
BO4Mob: Bayesian Optimization Benchmarks for High-Dimensional Urban Mobility Problem
Seunghee Ryu, Donghoon Kwon, Seongjin Choi +3
We introduce \textbf{BO4Mob}, a new benchmark framework for high-dimensional Bayesian Optimization (BO), driven by the challenge of origin-destination (OD) travel demand estimation…