7 papers
SkillAligner: Treating Retrieved Skills as Adaptable Drafts at Execution Time
Qinfeng Li, Dalin He, Yuntai Bao +7
General-purpose skills promise reusable procedural knowledge for language agents, yet semantic relevance does not guarantee execution utility: a retrieved skill may encode assumpti…
AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction
Qinfeng Li, Yuntai Bao, Xinyan Yu +8
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, co…
Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
Yuntai Bao, Qinfeng Li, Xinyan Yu +6
Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effec…
PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts
Qinfeng Li, Yuntai Bao, Jianghui Hu +5
LLM agents rely on prompts to implement task-specific capabilities based on foundation LLMs, making agent prompts valuable intellectual property. However, in untrusted deployments,…
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
Yuntai Bao, Xuhong Zhang, Jintao Chen +7
Intervention-based model steering offers a lightweight and interpretable alternative to prompting and fine-tuning. However, by adapting strong optimization objectives from fine-tun…
Scalable Multi-Stage Influence Function for Large Language Models via Eigenvalue-Corrected Kronecker-Factored Parameterization
Yuntai Bao, Xuhong Zhang, Tianyu Du +4
Pre-trained large language models (LLMs) are commonly fine-tuned to adapt to downstream tasks. Since the majority of knowledge is acquired during pre-training, attributing the pred…