activity
20242026
most citedAutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving

2 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.AI2026

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

Xiangning Lin, Shenzhe Zhu, Shu Yang +23

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are r…

cs.CL2026

Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing

Shenzhe Zhu, Haoqian Zhang, Xu Yang +7

Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that produced it. Final text alone cann…

cs.LG2026

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

Yi Nian, Tiankai Yang, Yudi Zhang +7

Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data selection methods typically scor…

cs.AI2026

When Only the Final Text Survives: Implicit Execution Tracing for Multi-Agent Auditing

Yi Nian, Haosen Cao, Shenzhe Zhu +4

When a multi-agent system produces an incorrect or harmful answer, who is accountable if execution logs and agent identifiers are unavailable? In practice, generated content is oft…

cs.AI2025

The Automated but Risky Game: Modeling and Benchmarking Agent-to-Agent Negotiations and Transactions in Consumer Markets

Shenzhe Zhu, Jiao Sun, Yi Nian +3

AI agents are increasingly used in consumer-facing applications to assist with tasks such as product search, negotiation, and transaction execution. In this paper, we explore a fut…

cs.CR2025

JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

Yi Nian, Shenzhe Zhu, Yuehan Qin +4

Multimodal large language models (MLLMs) excel in vision-language tasks but also pose significant risks of generating harmful content, particularly through jailbreak attacks. Jailb…