1 citations · 1 across the 2 of their papers we have counts for
5 papers
AI Security Leaderboard: Methodology, Results and Minimal Standard
Jasper Timm, Lukas Struppek, Ziwei Xu +12
The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FARAI Minimal Stan…
AI Agents Alone Are Not (Yet) Sufficient for Social Simulation
Yiming Li, Dacheng Tao
Recent advances in large language models (LLMs) have spurred growing interest in using LLM-integrated agents for social simulation, often under the implicit assumption that realist…
VIGIL: Towards Edge-Extended Agentic AI for Enterprise IT Support
Sarthak Ahuja, Neda Kordjazi, Evren Yortucboylu +7
Enterprise IT support is constrained by heterogeneous devices, evolving policies, and long-tail failure modes that are difficult to resolve centrally. We present VIGIL, an edge-ext…
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
Songze Li, Jiameng Cheng, Yiming Li +2
By integrating language understanding with perceptual modalities such as images, multimodal large language models (MLLMs) constitute a critical substrate for modern AI systems, par…
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
Xiaojun Jia, Jie Liao, Qi Guo +11
Recent advances in multi-modal large language models (MLLMs) have enabled unified perception-reasoning capabilities, yet these systems remain highly vulnerable to jailbreak attacks…