collaborators

6 papers

cs.CL2026

Models That Know How Evaluations Are Designed Score Safer

Katharina Deckenbach, Haritz Puerto, Jonas Geiping +1

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such a…

cs.CL2026

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

Haritz Puerto, Haonan Li, Xudong Han +2

Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit…

cs.CL2025

C-SEO Bench: Does Conversational SEO Work?

Haritz Puerto, Martin Gubri, Tommaso Green +2

Large Language Models (LLMs) are transforming search engines into Conversational Search Engines (CSE). Consequently, Search Engine Optimization (SEO) is being shifted into Conversa…

cs.CL2025

Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers

Tommaso Green, Martin Gubri, Haritz Puerto +2

We study privacy leakage in the reasoning traces of large reasoning models used as personal agents. Unlike final outputs, reasoning traces are often assumed to be internal and safe…

cs.CL2025

Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs

Haritz Puerto, Tilek Chubakov, Xiaodan Zhu +2

Requiring a large language model (LLM) to generate intermediary reasoning steps, known as Chain of Thought (CoT), has been shown to be an effective way of boosting performance. Pre…

cs.CL2025

Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models

Haritz Puerto, Martin Gubri, Sangdoo Yun +1

Membership inference attacks (MIA) attempt to verify the membership of a given data sample in the training set for a model. MIA has become relevant in recent years, following the r…