collaborators

6 papers

q-bio.BM2025

Steering Protein Language Models

Long-Kai Huang, Rongyi Zhu, Bing He +1

Protein Language Models (PLMs), pre-trained on extensive evolutionary data from natural proteins, have emerged as indispensable tools for protein design. While powerful, PLMs often…

cs.CL2025

Self-Improving Model Steering

Rongyi Zhu, Yuhui Wang, Tanqiu Jiang +2

Model steering represents a powerful technique that dynamically aligns large language models (LLMs) with human preferences during inference. However, conventional model-steering me…

cs.CV2025

Dynamic Token Reweighting for Robust Vision-Language Models

Tanqiu Jiang, Jiacheng Liang, Rongyi Zhu +3

Large vision-language models (VLMs) are highly vulnerable to multimodal jailbreak attacks that exploit visual-textual interactions to bypass safety guardrails. In this paper, we pr…

cs.LG2025

Self-Destructive Language Model

Yuhui Wang, Rongyi Zhu, Ting Wang

Harmful fine-tuning attacks pose a major threat to the security of large language models (LLMs), allowing adversaries to compromise safety guardrails with minimal harmful data. Whi…

cs.LG2025

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

Jiacheng Liang, Tanqiu Jiang, Yuhui Wang +3

This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its core, AutoRAN pioneers an execution…

cs.LG2025

GraphRAG under Fire

Jiacheng Liang, Yuhui Wang, Changjiang Li +4

GraphRAG advances retrieval-augmented generation (RAG) by structuring external knowledge as multi-scale knowledge graphs, enabling language models to integrate both broad context a…