From the 1 of 50 linked papers with an AI index.
1 citations · 2 across the 17 of their papers we have counts for
5 papers · 1 filter
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang +23
The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…
STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems
Guijia Zhang, Shu Yang, Xilin Gong +1
Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability surface before deployment, bu…
MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
Jingyu Hu, Shu Yang, Xilin Gong +3
Large Reasoning Models (LRMs) suffer from sycophantic behavior, where models tend to agree with users' incorrect beliefs and follow misinformation rather than maintain independent…
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
Liangliang You, Junchi Yao, Shu Yang +3
While multimodal large language models excel at various tasks, they still suffer from hallucinations, which limit their reliability and scalability for broader domain applications.…
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
Lijie Hu, Liang Liu, Shu Yang +5
Large Language Models have demonstrated remarkable abilities across various tasks, with Chain-of-Thought (CoT) prompting emerging as a key technique to enhance reasoning capabiliti…