activity
20242026
most citedMeasuring Agents in Production

3 citations · 3 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CY20263 cited

Measuring Agents in Production

Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo +22

LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first syst…

cs.AI2025

LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess

Sai Kolasani, Maxim Saplin, Nicholas Crispino +5

We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extend…

cs.CL2025

BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation

Alan Zhu, Parth Asawa, Jared Quincy Davis +5

As the demand for high-quality data in model training grows, researchers and developers are increasingly generating synthetic data to tune and train LLMs. However, current data gen…

cs.AI2025

Optimizing Model Selection for Compound AI Systems

Lingjiao Chen, Jared Quincy Davis, Boris Hanin +4

Compound AI systems that combine multiple LLM calls, such as self-refine and multi-agent-debate, achieve strong performance on many AI tasks. We address a core question in optimizi…

cs.SE2024

Specifications: The missing link to making the development of LLM systems an engineering discipline

Ion Stoica, Matei Zaharia, Joseph Gonzalez +8

Despite the significant strides made by generative AI in just a few short years, its future progress is constrained by the challenge of building modular and robust systems. This ca…