7 papers
Transformer-Based Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions
Vidur Sinha, Muhammed Ustaomeroglu, Guannan Qu
Multi-agent reinforcement learning (MARL) has shown promise for large-scale network control, yet existing methods face two major limitations. First, they typically rely on an expon…
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
Baris Askin, Muhammed Ustaomeroglu, Anupam Nayak +3
Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that e…
BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking
Muhammed Ustaomeroglu, Guannan Qu
Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior, yet also develops undesirable o…
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with -step Policy Gradients
Alex DeWeese, Guannan Qu
This work revisits standard policy gradient methods used on restricted policy classes, which are known to get stuck in suboptimal critical points. We identify an important cause fo…
Towards Effective Theory of LLMs: A Representation Learning Approach
Muhammed Ustaomeroglu, Guannan Qu
We propose Representational Effective Theory (RET), a framework for describing large language model computation in terms of learned macrostates rather than microscopic details. RET…
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
Muhammed Ustaomeroglu, Baris Askin, Gauri Joshi +2
The extent to which decoder-only language models (LMs) engage in planning, that is, organizing intermediate computations to support coherent long-range generation, remains an impor…