3 citations · 8 across the 15 of their papers we have counts for
15 papers
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
Yuxuan Zhu, Peng Pu
Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce Telemet…
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
Yuxuan Zhu, Rohan Alur, Daniel Kang
While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the result…
MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling
Zhong Li, Qi Huang, Yuxuan Zhu +6
Optimization modeling translates real decision-making problems into mathematical optimization models and solver-executable implementations. Although language models are increasingl…
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning
Ruijun Huang, Zhiqiao Kang, Yuxuan Zhu +5
The accurate extraction of scientific measurements from literature is a critical yet challenging task in AI4Science, enabling large-scale analysis and integration of quantitative r…
Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees
Yuxuan Zhu, Tengjun Jin, Chenghao Mo +1
Analytical join queries over unstructured data are increasingly prevalent in data analytics. Applying machine learning (ML) models to label every pair in the cross product of table…
Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering
Yuxuan Zhu, Tengjun Jin, Yoojin Choi +1
Translating natural language questions to SQL queries (Text-to-SQL) is a long-standing problem in database research. Recent efforts have focused on improving accuracy by building i…