Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Training Language Models via Neural Cellular Automata
Dan Lee, Seungwook Han, Akarsh Kumar +1
Pre-training is crucial for large language models (LLMs), as it is when most representations and capabilities are acquired. However, natural language pre-training has problems: hig…
cs.LG2024
Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs
Daniel J. Lee, Stefan Heimersheim
Sensitive directions experiments attempt to understand the computational features of Language Models (LMs) by measuring how much the next token prediction probabilities change by p…