1 citations · 1 across the 8 of their papers we have counts for
4 papers · 1 filter
Scaling Muon for Diffusion Transformers
Chenghao Li, Xiao Han, Xinxin Huang +22
The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion…
: Stratified Scaling Search for Test-Time in Diffusion Language Models
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +6
Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without additional training. However, n…
Continuous-Utility Direct Preference Optimization
Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal +6
Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasonin…
On the Fundamental Limits of LLMs at Scale
Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal +13
Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) re…