5 papers
Self-Improving Model Steering
Rongyi Zhu, Yuhui Wang, Tanqiu Jiang +2
Model steering represents a powerful technique that dynamically aligns large language models (LLMs) with human preferences during inference. However, conventional model-steering me…
Your Agent Can Defend Itself against Backdoor Attacks
Li Changjiang, Liang Jiacheng, Cao Bochuan +2
Despite their growing adoption across domains, large language model (LLM)-powered agents face significant security risks from backdoor attacks during training and fine-tuning. Thes…
GraphRAG under Fire
Jiacheng Liang, Yuhui Wang, Changjiang Li +4
GraphRAG advances retrieval-augmented generation (RAG) by structuring external knowledge as multi-scale knowledge graphs, enabling language models to integrate both broad context a…
CopyrightMeter: Revisiting Copyright Protection in Text-to-image Models
Naen Xu, Changjiang Li, Tianyu Du +8
Text-to-image diffusion models have emerged as powerful tools for generating high-quality images from textual descriptions. However, their increasing popularity has raised signific…
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
Jiacheng Liang, Zian Wang, Lauren Hong +2
Various watermarking methods (``watermarkers'') have been proposed to identify LLM-generated texts; yet, due to the lack of unified evaluation platforms, many critical questions re…