Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment
Fred Zhangzhi Peng, Alexis Fox, Anru R. Zhang +1
Diffusion language models (DLMs) have recently demonstrated capabilities that complement standard autoregressive (AR) models, particularly in non-sequential generation and bidirect…
cs.LG2026
Interpretable-by-Design Transformers via Architectural Stream Independence
Clayton Kerce, Alexis Fox
While transformers achieve strong performance, their internal decision-making processes remain opaque. We investigate whether architectural constraints can enforce interpretability…