Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PACE: A Proxy for Agentic Capability Evaluation
Yueqi Song, Lintang Sutawika, Jiarui Liu +8
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars…
cs.AI2026
ReasonOps: Operator Segmentation for LLM Reasoning Traces
Daniel Lee, Owen Queen, James Zou
Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structure. Previous methods develop…