From the 1 of 11 linked papers with an AI index.
11 papers
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification
Yihang Chen, Pin Qian, Su Wang +4
Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empir…
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
Chong Peng, Pin Qian, Su Wang +2
Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before us…
When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost
Pin Qian, Su Wang, Chong Peng +5
Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations oft…
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
Pin Qian, Su Wang, Yihang Chen +5
Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in…
Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened
Su Wang, Pin Qian, Yifan Lin +5
The paper investigates how self‑improving AI agents can hallucinate non‑existent failures and create unnecessary guardrails, introducing a deterministic Counterfactual Fabrication…
Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
Lifei Liu, Haoran Yu, Xiaochong Jiang +3
Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that…