Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
MiniMax, :, Aili Chen +125
We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combin…
cs.CL2024
CALICO: Conversational Agent Localization via Synthetic Data Generation
Andy Rosenbaum, Pegah Kharazmi, Ershad Banijamali +8
We present CALICO, a method to fine-tune Large Language Models (LLMs) to localize conversational agent training data from one language to another. For slots (named entities), CALIC…