8 papers
The Context-Ready Transformer
Mahesh Godavarti
We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes each token before it enters t…
Why Do Accumulated Transformations Extrapolate?
Mahesh Godavarti
PaTH Attention showed that replacing RoPE's position-indexed rotations with accumulated data-dependent Householder reflections yields strong length extrapolation, though performanc…
Convergence of Differential Entropies -- II
Mahesh Godavarti
We show that under convergence in measure of probability density functions, differential entropy converges whenever the entropy integrands are uniformly integrable…
Knowledge Graph and Hypergraph Transformers with Repository-Attention and Journey-Based Role Transport
Mahesh Godavarti
We present a concise architecture for joint training on sentences and structured data while keeping knowledge and language representations separable. The model treats knowledge gra…
JoFormer (Journey-based Transformer): Theory and Empirical Analysis on the Tiny Shakespeare Dataset
Mahesh Godavarti
Transformers have demonstrated remarkable success in sequence modeling, yet effectively incorporating positional information remains a challenging and active area of research. In t…
Directional Non-Commutative Monoidal Embeddings for MNIST
Mahesh Godavarti
We present an empirical validation of the directional non-commutative monoidal embedding framework recently introduced in prior work~\cite{Godavarti2025monoidal}. This framework de…