6 papers
SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment
Xixun Lin, Yang Liu, Yancheng Chen +8
The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context management, and state persist…
WorldCup Sampling for Multi-bit LLM Watermarking
Yidan Wang, Yubing Ren, Yanan Cao +1
As large language models (LLMs) generate increasingly human-like text, watermarking has emerged as a promising solution for reliable attribution beyond mere detection. While multi-…
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
Hao Li, Yubing Ren, Yanan Cao +4
With the rapid development of cloud-based services, large language models have become increasingly accessible through various web platforms. However, this accessibility has also le…
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
Zhuoshang Wang, Yubing Ren, Yanan Cao +3
While watermarking serves as a critical mechanism for LLM provenance, existing secret-key schemes tightly couple detection with injection, requiring access to keys or provider-side…
Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text Detection
Xiaowei Zhu, Yubing Ren, Fang Fang +3
The rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation diss…
LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
Xixun Lin, Yucheng Ning, Jingwen Zhang +21
Driven by the rapid advancements of Large Language Models (LLMs), LLM-based agents have emerged as powerful intelligent systems capable of human-like cognition, reasoning, and inte…