Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs
Debopam Sanyal, Anantharaman Iyer, Alind Khare +5
Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent work demonstrates that stitching…
cs.LG2026
Engineering Verifiable Modularity in Transformers via Per-Layer Supervision
J. Clayton Kerce
Transformers resist surgical control. Ablating an attention head identified as critical for capitalization produces minimal behavioral change because distributed redundancy compens…
cs.LG2026
Interpretable-by-Design Transformers via Architectural Stream Independence
Clayton Kerce, Alexis Fox
While transformers achieve strong performance, their internal decision-making processes remain opaque. We investigate whether architectural constraints can enforce interpretability…