10 papers
Memory-Induced Tool-Drift in LLM Agents
Mahavir Dabas, Jihyun Jeong, Ming Jin +1
Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinning contemporary production sy…
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
Ahmad Al-Tawaha, Shangding Gu, Peizhi Niu +2
Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under adversarial conditions such…
Optimizing Product Provenance Verification using Data Valuation Methods
Raquib Bin Yousuf, Hoang Anh Just, Shengzhe Xu +8
Determining and verifying product provenance remains a critical challenge in global supply chains, particularly as geopolitical conflicts and shifting borders create new incentives…
DiPT: Enhancing LLM reasoning through diversified perspective-taking
Hoang Anh Just, Mahavir Dabas, Lifu Huang +2
Existing work on improving language model reasoning typically explores a single solution path, which can be prone to errors. Inspired by perspective-taking in social studies, this…
Data-Centric Human Preference with Rationales for Direct Preference Alignment
Hoang Anh Just, Ming Jin, Anit Sahu +2
Aligning language models with human preferences through reinforcement learning from human feedback is crucial for their safe and effective deployment. The human preference is typic…
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Mahavir Dabas, Si Chen, Charles Fleming +2
Safety alignment is crucial for large language models (LLMs) to resist malicious instructions but often results in over-refusals, where benign prompts are unnecessarily rejected, i…