From the 1 of 15 linked papers with an AI index.
15 papers
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
Yuxuan Zhu, Peng Pu
Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce Telemet…
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
Yuxuan Zhu, Rohan Alur, Daniel Kang
The paper derives the first non‑vacuous PAC‑Bayes generalization bounds for parameter‑efficient reinforcement learning with verifiable rewards applied to billion‑parameter language…
Measuring Agents in Production
Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo +22
LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first syst…
MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling
Zhong Li, Qi Huang, Yuxuan Zhu +6
Optimization modeling translates real decision-making problems into mathematical optimization models and solver-executable implementations. Although language models are increasingl…
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning
Ruijun Huang, Zhiqiao Kang, Yuxuan Zhu +5
The accurate extraction of scientific measurements from literature is a critical yet challenging task in AI4Science, enabling large-scale analysis and integration of quantitative r…
Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering
Yuxuan Zhu, Tengjun Jin, Yoojin Choi +1
Translating natural language questions to SQL queries (Text-to-SQL) is a long-standing problem in database research. Recent efforts have focused on improving accuracy by building i…