activity
20242026
most citedJailbreak Attacks and Defenses Against Large Language Models: A Survey

7 citations · 9 across the 9 of their papers we have counts for

collaborators

17 papers

cs.LG2026

CHASM: Unveiling Covert Advertisements on Chinese Social Media

Jingyi Zheng, Tianyi Hu, Yule Liu +5

Current benchmarks for evaluating large language models (LLMs) in social media moderation completely overlook a serious threat: covert advertisements, which disguise themselves as…

cs.AI2026

Chain-of-Authorization: Embedding authorization into large language models

Yang Li, Yule Liu, Xinlei He +3

Although Large Language Models (LLMs) have evolved from text generators into the cognitive core of modern AI systems, their inherent lack of authorization awareness exposes these s…

cs.CR2025

"To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios

Zhen Sun, Zongmin Zhang, Deqi Liang +8

As LLMs become more common, non-expert users can pose risks, prompting extensive research into jailbreak attacks. However, most existing black-box jailbreak attacks rely on hand-cr…

cs.CL20251 cited

ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities

Wenhan Dong, Zhen Sun, Yuemeng Zhao +9

Large language models (LLMs) have demonstrated potential in educational applications, yet their capacity to accurately assess the cognitive alignment of reading materials with stud…

cs.AI2025

GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts

Jingyi Zheng, Zifan Peng, Yule Liu +4

Smart contracts are trustworthy, immutable, and automatically executed programs on the blockchain. Their execution requires the Gas mechanism to ensure efficiency and fairness. How…

cs.CL2025

Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks

Wenhan Dong, Tianyi Hu, Jingyi Zheng +5

Multi-round incomplete information tasks are crucial for evaluating the lateral thinking capabilities of large language models (LLMs). Currently, research primarily relies on multi…