3 papers
cs.LG2026
The Depth Delusion: Why Transformers Should Be Wider, Not Deeper
Md Muhtasim Munif Fahim, Md Rezaul Karim
Neural scaling laws describe how language model loss decreases with parameters and data, but treat architecture as interchangeable--a billion parameters could arise from a shallow-…
cs.LG2026
Pre-trained Encoders for Global Child Development: Transfer Learning Enables Deployment in Data-Scarce Settings
Md Muhtasim Munif Fahim, Md Rezaul Karim
A large number of children experience preventable developmental delays each year, yet the deployment of machine learning in new countries has been stymied by a data bottleneck: rel…
cs.LG2026
The Dependency Divide: An Interpretable Machine Learning Framework for Profiling Student Digital Satisfaction in the Bangladesh Context
Md Muhtasim Munif Fahim, Humyra Ankona, Md Monimul Huq +1
Background: While digital access has expanded rapidly in resource-constrained contexts, satisfaction with digital learning platforms varies significantly among students with seemin…