collaborators

7 papers

cs.LG2026

Diffuse AI Control on Fuzzy Tasks

Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar +1

AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned wit…

cs.LG2026

Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols

Mikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou +4

AI control protocols serve as a defense mechanism to stop untrusted LLM agents from causing harm in autonomous settings. Prior work treats this as a security problem, stress testin…

cs.AI2026

Control Tax: The Price of Keeping AI in Check

Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre +1

The rapid integration of agentic AI into high-stakes real-world applications requires robust oversight mechanisms. The emerging field of AI Control (AIC) aims to provide such an ov…

cs.LG2025

One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models

Viacheslav Surkov, Chris Wendler, Antonio Mari +5

For large language models (LLMs), sparse autoencoders (SAEs) have been shown to decompose intermediate representations that often are not interpretable directly into sparse sums of…

quant-ph2025

Learning pure quantum states (almost) without regret

Josep Lumbreras, Mikhail Terekhov, Marco Tomamichel

We initiate the study of sample-optimal quantum state tomography with minimal disturbance to the samples. Can we efficiently learn a precise description of a quantum state through…

q-fin.PM2025

Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards

Daniil Karzanov, Rubén Garzón, Mikhail Terekhov +3

This paper introduces a novel agent-based approach for enhancing existing portfolio strategies using Proximal Policy Optimization (PPO). Rather than focusing solely on traditional…