1 citations · 3 across the 13 of their papers we have counts for
10 papers · 1 filter
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
Yutao Mou, Zhangchi Xue, Lijun Li +4
While LLM-based agents can interact with environments via invoking external tools, their expanded capabilities also amplify security risks. Monitoring step-level tool invocation be…
Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG
Bo Li, Tian Tian, Zhenghua Xu +3
Dynamic retrieval-augmented generation (RAG) allows large language models (LLMs) to fetch external knowledge on demand, offering greater adaptability than static RAG. A central cha…
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
Yutao Mou, Xiaoling Zhou, Yuxiao Luo +2
Safety alignment is essential for building trustworthy artificial intelligence, yet it remains challenging to enhance model safety without degrading general performance. Current ap…
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
Muxi Diao, Yutao Mou, Keqing He +6
The safety of Large Language Models (LLMs) is crucial for the development of trustworthy AI applications. Existing red teaming methods often rely on seed instructions, which limits…
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
Zhuohao Yu, Xingru Jiang, Weizheng Gu +4
Watermarking LLM-generated text is critical for content attribution and misinformation prevention. However, existing methods compromise text quality, require white-box model access…
MPL: Multiple Programming Languages with Large Language Models for Information Extraction
Bo Li, Gexiang Fang, Wei Ye +4
Recent research in information extraction (IE) focuses on utilizing code-style inputs to enhance structured output generation. The intuition behind this is that the programming lan…