activity
20242026
most citedPandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

1 citations · 1 across the 7 of their papers we have counts for

collaborators

7 papers

cs.AI2026

CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models

Haibo Tong, Zeyang Yue, Feifei Zhao +6

Whether Large Language Models (LLMs) truly possess human-like Theory of Mind (ToM) capabilities has garnered increasing attention. However, existing benchmarks remain largely restr…

cs.AI2025

Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense

Guobin Shen, Dongcheng Zhao, Haibo Tong +3

Ensuring Large Language Model (LLM) safety remains challenging due to the absence of universal standards and reliable content validators, making it difficult to obtain effective tr…

cs.CR2025

Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks

Haibo Tong, Dongcheng Zhao, Guobin Shen +4

The remarkable capabilities of Large Language Models (LLMs) have raised significant safety concerns, particularly regarding "jailbreak" attacks that exploit adversarial prompts to…

cs.CR20251 cited

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Guobin Shen, Dongcheng Zhao, Linghao Feng +8

Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit ha…

cs.AI2025

Super Co-alignment of Human and AI for Sustainable Symbiotic Society

Yi Zeng, Feifei Zhao, Yuwei Wang +13

As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, d…

cs.AI2025

Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind

Haibo Tong, Enmeng Lu, Yinqian Sun +4

With the widespread application of Artificial Intelligence (AI) in human society, enabling AI to autonomously align with human values has become a pressing issue to ensure its sust…