3 papers
cs.LG2026
Localizing RL-Induced Tool Use to a Single Crosscoder Feature
Andrii Shportko, Shubham Bhokare, Ahmed Zeyad A Alzahrani +3
Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poo…
cs.LG2026
Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy
Andrii Shportko
Large language models can rewrite text to embed hidden payloads while preserving surface-level meaning, a capability that opens covert channels between cooperating AI systems and p…
cs.LG2025
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
Matthew Khoriaty, Andrii Shportko, Gustavo Mercier +1
Recent developments in Large Language Model (LLM) capabilities have brought great potential but also posed new risks. For example, LLMs with knowledge of bioweapons, advanced chemi…