5 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
Maximal Update Parametrization and Zero-Shot Hyperparameter Transfer for Fourier Neural Operators
Shanda Li, Shinjae Yoo, Yiming Yang
Fourier Neural Operators (FNOs) offer a principled approach for solving complex partial differential equations (PDEs). However, scaling them to handle more complex PDEs requires in…
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
Baihe Huang, Shanda Li, Tianhao Wu +5
Test-time scaling paradigms have significantly advanced the capabilities of large language models (LLMs) on complex tasks. Despite their empirical success, theoretical understandin…
CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
Weiwei Sun, Shengyu Feng, Shanda Li +1
Although LLM-based agents have attracted significant attention in domains such as software engineering and machine learning research, their role in advancing combinatorial optimiza…
TFG-Flow: Training-free Guidance in Multimodal Generative Flow
Haowei Lin, Shanda Li, Haotian Ye +4
Given an unconditional generative model and a predictor for a target property (e.g., a classifier), the goal of training-free guidance is to generate samples with desirable target…