works on

From the 1 of 27 linked papers with an AI index.

collaborators

27 papers

cs.AI2026

Measuring Activation Control in Large Language Models

Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar +1

Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware model…

cs.AI2026

Item Response Theory for AI Safety

Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1

Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because b…

cs.AI2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…

cs.LG2026

Detecting CSAM Text-to-Image LoRAs From Weights

David Demitri Africa, Cate Heine, Nadine Staes-Polet +1

Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abus…

cs.AI2026

Persona Cartography: Charting Language Model Personality Traits in Weight Space

Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk +4

Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and control…

cs.CR2026

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

Jimmy Laurence Rippin, Simon C. Marshall, David Demitri Africa +1

Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defence to these collusion attempts…