Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aakshita Chandiramani +544
We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemo…
cs.LG2025
Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
Peter Belcak, Greg Heinrich, Jan Kautz +1
Finetuning language models for a new domain inevitably leads to the deterioration of their general performance. This becomes more pronounced the more limited the finetuning data re…
cs.LG2024
Tiny Transformers Excel at Sentence Compression
Peter Belcak, Roger Wattenhofer
It is staggering that words of the English language, which are on average represented by 5--6 bytes of ASCII, require as much as 24 kilobytes when served to large language models.…