Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
VFUSE: Virulent Feature Understanding with Sparse autoEncoders
Michael Yu, Matthew L. Olson
Generative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, w…
cs.LG2026
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
John Yan, Michael Yu, Yuqi Sun +3
Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over train…