most citedMeasuring Agents in Production

3 citations · 8 across the 15 of their papers we have counts for

collaborators

15 papers

cs.AI2026

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

Yuxuan Zhu, Peng Pu

Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce Telemet…

cs.LG2026

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

Yuxuan Zhu, Rohan Alur, Daniel Kang

While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the result…

cs.AI2026

MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling

Zhong Li, Qi Huang, Yuxuan Zhu +6

Optimization modeling translates real decision-making problems into mathematical optimization models and solver-executable implementations. Although language models are increasingl…

cs.CL2026

MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning

Ruijun Huang, Zhiqiao Kang, Yuxuan Zhu +5

The accurate extraction of scientific measurements from literature is a critical yet challenging task in AI4Science, enabling large-scale analysis and integration of quantitative r…

cs.DB2026

Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees

Yuxuan Zhu, Tengjun Jin, Chenghao Mo +1

Analytical join queries over unstructured data are increasingly prevalent in data analytics. Applying machine learning (ML) models to label every pair in the cross product of table…

cs.DB2026

Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering

Yuxuan Zhu, Tengjun Jin, Yoojin Choi +1

Translating natural language questions to SQL queries (Text-to-SQL) is a long-standing problem in database research. Recent efforts have focused on improving accuracy by building i…