6 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Apple Intelligence Foundation Language Models: Tech Report 2025
Ethan Li, Anders Boesen Lindbo Larsen, Chen Zhang +395
We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model opti…
cs.LG2023★ 6 cited
Toward Understanding Why Adam Converges Faster Than SGD for Transformers
Yan Pan, Yuanzhi Li
While stochastic gradient descent (SGD) is still the most popular optimization algorithm in deep learning, adaptive algorithms such as Adam have established empirical advantages ov…