5 papers
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
Phillip Howard, Xin Su, Allen Roush +2
High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques have helped mitigate drawbacks…
Dynamic Latent Routing
Fangyuan Yu, Xin Su, Amir Abdullah
We investigate the temporal concatenation of sub-policies in Markov Decision Processes (MDP) with time-varying reward functions. We introduce General Dijkstra Search (GDS), and pro…
Spectral Superposition: A Theory of Feature Geometry
Georgi Ivanov, Narmeen Oozeer, Shivam Raval +3
Neural networks represent more features than they have dimensions via superposition, forcing features to share representational space. Current methods decompose activations into sp…
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
Philip Quirke, Narmeen Oozeer, Chaithanya Bandi +8
This position paper argues that the prevailing trajectory toward ever larger, more expensive generalist foundation models controlled by a handful of companies limits innovation and…
Interpreting Learned Feedback Patterns in Large Language Models
Luke Marks, Amir Abdullah, Clement Neo +4
Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferen…