collaborators

12 papers

cs.CL2026

Noise Floor Audit for Agent Benchmarks

Yihang Chen, Pin Qian, Su Wang +4

We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At tempera…

cs.AI2026

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

Yihang Chen, Pin Qian, Su Wang +4

Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empir…

cs.AI2026

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

Chong Peng, Pin Qian, Su Wang +2

Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before us…

cs.LG2026

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

Pin Qian, Su Wang, Chong Peng +5

Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations oft…

cs.LG2026

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

Pin Qian, Su Wang, Yihang Chen +5

Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in…

cs.CR2026

Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened

Su Wang, Pin Qian, Yifan Lin +5

Self-improving AI agents are designed to learn from their mistakes. We show they can also hallucinate mistakes that never happened. We study this failure mode in automated harness…