145 citations · 145 across the 3 of their papers we have counts for
3 papers
cs.DC2025
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
Lingling Zeng, Gen Zhang, Jialin Peng +3
As AI cluster sizes continue to expand and the demand for large-language-model (LLM) training and inference workloads grows rapidly, traditional scheduling systems face significant…
cs.CL2025
Token Masking Improves Transformer-Based Text Classification
Xianglong Xu, John Bowen, Rojin Taheri
While transformer-based models achieve strong performance on text classification, we explore whether masking input tokens can further enhance their effectiveness. We propose token…
cs.CL2024★ 145 cited
Gemma 2: Improving Open Language Models at a Practical Size
Gemma Team, Morgane Riviere, Shreya Pathak +195
In this work, we introduce Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters. In th…