From the 1 of 54 linked papers with an AI index.
6 citations · 17 across the 26 of their papers we have counts for
6 papers · 1 filter
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
Marcus J. Min, Mike He, Zhaoyu Li +5
The paper proposes shifting autoformalization from isolated statements to theory-level, aiming to automatically translate whole bodies of mathematical knowledge—including axioms, d…
Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning
Wenwen Si, Insup Lee, Osbert Bastani
Inference-time computation has greatly enhanced the performance of large language models (LLMs) on challenging reasoning tasks, but this strategy can incur high inference costs. On…
Are AI Capabilities Increasing Exponentially? A Competing Hypothesis
Haosen Ge, Hamsa Bastani, Osbert Bastani
Rapidly increasing AI capabilities have substantial real-world consequences, ranging from AI safety concerns to labor market consequences. The Model Evaluation & Threat Research (M…
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
Sagnik Anupam, Davis Brown, Shuo Li +3
LLM web agents now browse and take actions on the open web, yet current agent evaluations are constrained to sandboxed environments or artificial tasks. We introduce BrowserArena,…
Effective Reinforcement Learning for Reasoning in Language Models
Lianghuan Huang, Shuo Li, Sagnik Anupam +2
Reinforcement learning (RL) has emerged as a promising strategy for improving the reasoning capabilities of language models (LMs) in domains such as mathematics and coding. However…
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Xinmeng Huang, Shuo Li, Edgar Dobriban +3
The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safe…