2 papers
cs.SE2025
State-of-the-art Small Language Coder Model: Mify-Coder
Abhinav Parmar, Abhisek Panigrahi, Abhishek Kumar Dwivedi +93
We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable a…
cs.LG2025
Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
Samarth Gupta, Raghudeep Gadde, Rui Chen +1
We challenge a fundamental assumption of diffusion models, namely, that a large number of latent-states or time-steps is required for training so that the reverse generative proces…