4 papers
Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
Haibo Jin, Suijin Wang, Xucheng Yu +2
Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two…
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
Xucheng Yu, Emily Knox, Haohan Wang
As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We s…
SCI-Defense: Defending Manipulation Attacks from Generative Engine Optimization
Xucheng Yu, Haibo Jin, Huimin Zeng +1
LLM-based ranking systems are vulnerable to Generative Engine Optimization (GEO) attacks, where adversaries inject semantic signals into product descriptions to artificially boost…
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
Han Wang, Yifan Sun, Brian Ko +8
Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer…