collaborators

19 papers

cs.SE2026

Fantastic Adaptive Taxonomies and How to Use Them

Mert Cemri, Andrei Cojocaru, Melissa Pan +9

An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory selection, prompt and workflow optimiza…

cs.AI2026

LLM-as-a-Verifier: A General-Purpose Verification Framework

Jacky Kwok, Shulu Li, Pranav Atreya +6

Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the abi…

cs.LG2026

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Aaron J. Li, Hao Huang, Youngmin Park +6

Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preference data that better r…

cs.RO2026

Playful Agentic Robot Learning

Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reu…

cs.AI2026

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Xirui Li, Ming Li, Ion Stoica +2

Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that what is needed is not just a dat…

cs.CY2026

Measuring Agents in Production

Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo +22

LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first syst…