8 papers
Motif-Video 2B: Technical Report
Junghwan Lim, Wai Ting Cheung, Minsu Ha +25
Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video qualit…
Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models
Junhyeok Lee, Kyu Sung Choi
Mixture-of-Experts (MoE) language models are universally sensitive to demographic content at the routing level, yet exploiting this sensitivity for fairness control is structurally…
citecheck: An MCP Server for Automated Bibliographic Verification and Repair in Scholarly Manuscripts
Junhyeok Lee
Reference lists in scholarly manuscripts frequently contain errors, including incorrect identifiers, incomplete metadata, misattributed authors, and mismatches between preprint and…
Motif-2-12.7B-Reasoning: A Practitioner's Guide to RL Training Recipes
Junghwan Lim, Sungmin Lee, Dongseok Kim +23
We introduce Motif-2-12.7B-Reasoning, a 12.7B parameter language model designed to bridge the gap between open-weight systems and proprietary frontier models in complex reasoning a…
Motif 2 12.7B technical report
Junghwan Lim, Sungmin Lee, Dongseok Kim +22
We introduce Motif-2-12.7B, a new open-weight foundation model that pushes the efficiency frontier of large language models by combining architectural innovation with system-level…
Grouped Differential Attention
Junghwan Lim, Sungmin Lee, Dongseok Kim +7
The self-attention mechanism, while foundational to modern Transformer architectures, suffers from a critical inefficiency: it frequently allocates substantial attention to redunda…