5 papers
KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs
Debopam Sanyal, Anantharaman Iyer, Alind Khare +5
Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent work demonstrates that stitching…
Engineering Verifiable Modularity in Transformers via Per-Layer Supervision
J. Clayton Kerce
Transformers resist surgical control. Ablating an attention head identified as critical for capitalization produces minimal behavioral change because distributed redundancy compens…
Interpretable-by-Design Transformers via Architectural Stream Independence
Clayton Kerce, Alexis Fox
While transformers achieve strong performance, their internal decision-making processes remain opaque. We investigate whether architectural constraints can enforce interpretability…
The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
J. Clayton Kerce, Alexis Fox
Standard transformers entangle all computation in a single residual stream, obscuring which components perform which functions. We introduce the Dual-Stream Transformer, which deco…
GLIDR: Graph-Like Inductive Logic Programming with Differentiable Reasoning
Blair Johnson, Clayton Kerce, Faramarz Fekri
Differentiable inductive logic programming (ILP) techniques have proven effective at finding approximate rule-based solutions to link prediction and node classification problems on…