2 papers
cs.LG2026
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
Tobias Falke, Nicolas Anastassacos, Samson Tan +6
Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges due to routing complexity. Fu…
cs.CL2024
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…