3 papers
cs.LG2026
Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations
Pratyush Acharya, Nuraj Rimal, Habish Dhakal
We test whether the causal inner product of \citet{park2024linear} -- defined by the unembedding covariance -- enables cross-lingual concept transport. Across 17 models and 4…
cs.CY2026
Assessing the Pedagogical Readiness of Large Language Models as AI Tutors in Low-Resource Contexts: A Case Study of Nepal's K-10 Curriculum
Pratyush Acharya, Prasansha Bharati, Yokibha Chapagain +2
The integration of Large Language Models (LLMs) into educational ecosystems promises to democratize access to personalized tutoring, yet the readiness of these systems for deployme…
cs.LG2026
Grokking as a Variance-Limited Phase Transition: Spectral Gating and the Epsilon-Stability Threshold
Pratyush Acharya, Habish Dhakal
Standard optimization theories struggle to explain grokking, where generalization occurs long after training convergence. While geometric studies attribute this to slow drift, they…