Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
Weitao Feng, Lixu Wang, Peizhuo Lv +5
As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studies assume that attackers rely on superv…
cs.LG2026
How Implicit Bias Accumulates and Propagates in LLM Long-term Memory
Yiming Ma, Lixu Wang, Lionel Z. Wang +6
Long-term memory mechanisms enable Large Language Models (LLMs) to maintain continuity and personalization across extended interaction lifecycles, but they also introduce new and u…