activity
20242026
collaborators

9 papers

cs.CR2026

AI Security Leaderboard: Methodology, Results and Minimal Standard

Jasper Timm, Lukas Struppek, Ziwei Xu +12

The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FARAI Minimal Stan…

cs.AI2026

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

Yang Feng, Ziwei Xu, Xia Hu +1

Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques. However, fixed topologies often introduce persistent positional biases…

cs.CY2026

Buy versus Build an LLM: A Decision Framework for Governments

Jiahao Lu, Ziwei Xu, William Tjhi +4

Large Language Models (LLMs) represent a new frontier of digital infrastructure that can support a wide range of public-sector applications, from general purpose citizen services t…

cs.CL2026

Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention

Himanshu Singh, Ziwei Xu, A. V. Subramanyam +1

Large Language Models (LLMs) are powerful text generators, yet they can produce toxic or harmful content even when given seemingly harmless prompts. This presents a serious safety…

cs.CR2026

LLMs Can Unlearn Refusal with Only 1,000 Benign Samples

Yangyang Guo, Ziwei Xu, Si Liu +2

This study reveals a previously unexplored vulnerability in the safety alignment of Large Language Models (LLMs). Existing aligned LLMs predominantly respond to unsafe queries with…

cs.CL2025

Reasoning LLMs are Wandering Solution Explorers

Jiahao Lu, Ziwei Xu, Mohan Kankanhalli

Large Language Models (LLMs) have demonstrated impressive reasoning abilities through test-time computation (TTC) techniques such as chain-of-thought prompting and tree-based reaso…