works on

From the 1 of 54 linked papers with an AI index.

activity
20242026
most citedOptimal Program Synthesis via Abstract Interpretation

6 citations · 17 across the 26 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases

Marcus J. Min, Mike He, Zhaoyu Li +5

The paper proposes shifting autoformalization from isolated statements to theory-level, aiming to automatically translate whole bodies of mathematical knowledge—including axioms, d…

cs.AI2026

Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning

Wenwen Si, Insup Lee, Osbert Bastani

Inference-time computation has greatly enhanced the performance of large language models (LLMs) on challenging reasoning tasks, but this strategy can incur high inference costs. On…

cs.AI2026

Are AI Capabilities Increasing Exponentially? A Competing Hypothesis

Haosen Ge, Hamsa Bastani, Osbert Bastani

Rapidly increasing AI capabilities have substantial real-world consequences, ranging from AI safety concerns to labor market consequences. The Model Evaluation & Threat Research (M…

cs.AI2025

BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks

Sagnik Anupam, Davis Brown, Shuo Li +3

LLM web agents now browse and take actions on the open web, yet current agent evaluations are constrained to sandboxed environments or artificial tasks. We introduce BrowserArena,…

cs.AI2025

Effective Reinforcement Learning for Reasoning in Language Models

Lianghuan Huang, Shuo Li, Sagnik Anupam +2

Reinforcement learning (RL) has emerged as a promising strategy for improving the reasoning capabilities of language models (LMs) in domains such as mathematics and coding. However…

cs.AI2024

One-Shot Safety Alignment for Large Language Models via Optimal Dualization

Xinmeng Huang, Shuo Li, Edgar Dobriban +3

The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safe…