collaborators

18 papers

cs.CR2026

AI Security Leaderboard: Methodology, Results and Minimal Standard

Jasper Timm, Lukas Struppek, Ziwei Xu +12

The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FARAI Minimal Stan…

cs.AI2026

Large language models can effectively convince people to believe conspiracies

Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6

The paper investigates whether large language models can be used to persuade people to adopt or reject conspiracy beliefs, finding that LLMs can both increase and decrease belief d…

cs.SI2026

CrediBench: Building Web-Scale Network Datasets for Information Integrity

Emma Kondrup, Sebastian Sabry, Hussein Abdallah +9

Automatically assessing the credibility of online sources presents an invaluable tool for navigating today's information ecosystem. However, existing approaches either depend on sc…

cs.CR2026

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

Saad Hossain, Tom Tseng, Punya Syon Pandey +8

As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, be…

cs.CL2026

Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards

Punya Syon Pandey, Samuel Simko, Kellin Pelrine +1

As large language models (LLMs) gain popularity, their vulnerability to adversarial attacks emerges as a primary concern. While fine-tuning models on domain-specific datasets is of…

cs.AI2026

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Matthew Kowal, Goncalo Paulo, Louis Jaburi +6

As large language models are increasingly trained and fine-tuned, practitioners need methods to identify which training data drive specific behaviors, particularly unintended ones.…