2 papers
cs.SE2025
State-of-the-art Small Language Coder Model: Mify-Coder
Abhinav Parmar, Abhisek Panigrahi, Abhishek Kumar Dwivedi +93
We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable a…
cs.LG2025
Conformal Transformations for Symmetric Power Transformers
Saurabh Kumar, Jacob Buckman, Carles Gelada +1
Transformers with linear attention offer significant computational advantages over softmax-based transformers but often suffer from degraded performance. The symmetric power (sympo…