2 papers
cs.CR2026
Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation
Minh Tran, Cuong Dang, Tuc Nguyen +10
Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. At the same time,…
cs.LG2026
Adversarial Robustness of Activation Steering in Large Language Models
Kien Le, Thai Le
Activation steering has become a popular training-free method to control LLM behavior by injecting precomputed direction vectors into the model's residual stream at inference time.…