6 papers
Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models
Junxiang You, Junkai Chen, Yuhao He +3
Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent…
When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost
Pin Qian, Su Wang, Chong Peng +5
Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations oft…
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
Pin Qian, Su Wang, Yihang Chen +5
Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in…
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
Su Wang, Pin Qian, Yihang Chen +6
LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agentic AI systems: whether indivi…
Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG
Pin Qian, Su Wang, Xiaoyuan Wang +7
Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording. We study this diagnos…
Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning
Junkai Chen, Yuhao He, Junxiang You +3
Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive or restricted knowledge, raisin…