collaborators

6 papers

cs.CY2026

Agentic Safety is an Epistemic Property, Not a Behavioral One

Charles L. Wang, Keir Dorchen, Peter Jin

Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they prima…

cs.AI2026

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?

Tu Trinh, Mohamed Elfeki, Guangze Luo +9

Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete or ambiguous. The bottleneck is not raw capability, but judgm…

cs.AI2026

On The Statistical Limits of Self-Improving Agents

Charles L. Wang, Keir Dorchen, Peter Jin

We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: und…

cs.AI2025

MathBode: Measuring the Stability of LLM Reasoning using Frequency Response

Charles L. Wang

This paper presents MathBode, a dynamic diagnostic for mathematical reasoning in large language models (LLMs). Instead of one-shot accuracy, MathBode treats each parametric problem…

cs.AI2025

MI9: An Integrated Runtime Governance Framework for Agentic AI

Charles L. Wang, Trisha Singhal, Ameya Kelkar +1

Agentic AI systems capable of reasoning, planning, and executing actions present fundamentally distinct governance challenges compared to traditional AI models. Unlike conventional…

cs.CV2025

Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning

Ang Li, Charles Wang, Deqing Fu +9

Humans often use visual aids, for example diagrams or sketches, when solving complex problems. Training multimodal models to do the same, known as Visual Chain of Thought (Visual C…