4 papers
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
Tobias Falke, Nicolas Anastassacos, Samson Tan +6
Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges due to routing complexity. Fu…
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
Moulik Choraria, Xinbo Wu, Akhil Bhimaraju +5
Hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs, underlining a growing need for more efficient finetuning a…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
Semantically Grounded QFormer for Efficient Vision Language Understanding
Moulik Choraria, Xinbo Wu, Sourya Basu +5
General purpose Vision Language Models (VLMs) have received tremendous interest in recent years, owing to their ability to learn rich vision-language correlations as well as their…