3 papers
cs.LG2026
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
Pin Qian, Su Wang, Yihang Chen +5
Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in…
cs.SE2026
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
Su Wang, Pin Qian, Yihang Chen +6
LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agentic AI systems: whether indivi…
cs.AI2026
Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG
Pin Qian, Su Wang, Xiaoyuan Wang +7
Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording. We study this diagnos…