collaborators

5 papers

cs.LG2026

KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs

Debopam Sanyal, Anantharaman Iyer, Alind Khare +5

Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent work demonstrates that stitching…

cs.LG2026

Engineering Verifiable Modularity in Transformers via Per-Layer Supervision

J. Clayton Kerce

Transformers resist surgical control. Ablating an attention head identified as critical for capitalization produces minimal behavioral change because distributed redundancy compens…

cs.LG2026

Interpretable-by-Design Transformers via Architectural Stream Independence

Clayton Kerce, Alexis Fox

While transformers achieve strong performance, their internal decision-making processes remain opaque. We investigate whether architectural constraints can enforce interpretability…

cs.CL2026

The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling

J. Clayton Kerce, Alexis Fox

Standard transformers entangle all computation in a single residual stream, obscuring which components perform which functions. We introduce the Dual-Stream Transformer, which deco…

cs.AI2025

GLIDR: Graph-Like Inductive Logic Programming with Differentiable Reasoning

Blair Johnson, Clayton Kerce, Faramarz Fekri

Differentiable inductive logic programming (ILP) techniques have proven effective at finding approximate rule-based solutions to link prediction and node classification problems on…