3 papers
cs.CL2025
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
Ferdinand Kapl, Emmanouil Angelis, Tobias Höppe +4
Gradually growing the depth of Transformers during training can not only reduce training cost but also lead to improved reasoning performance, as shown by MIDAS (Saunshi et al., 20…
cs.CV2025
Breaking the Likelihood-Quality Trade-off in Diffusion Models by Merging Pretrained Experts
Yasin Esfandiari, Stefan Bauer, Sebastian U. Stich +1
Diffusion models for image generation often exhibit a trade-off between perceptual sample quality and data likelihood: training objectives emphasizing high-noise denoising steps yi…
cs.LG2025
Jasmine: A Simple, Performant and Scalable JAX-based World Modeling Codebase
Mihir Mahajan, Alfred Nguyen, Franz Srambical +1
While world models are increasingly positioned as a pathway to overcoming data scarcity in domains such as robotics, open training infrastructure for world modeling remains nascent…