2 citations · 4 across the 11 of their papers we have counts for
8 papers · 1 filter
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aakshita Chandiramani +544
We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemo…
Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
Peter Belcak, Greg Heinrich, Jan Kautz +1
Finetuning language models for a new domain inevitably leads to the deterioration of their general performance. This becomes more pronounced the more limited the finetuning data re…
Tiny Transformers Excel at Sentence Compression
Peter Belcak, Roger Wattenhofer
It is staggering that words of the English language, which are on average represented by 5--6 bytes of ASCII, require as much as 24 kilobytes when served to large language models.…
Fast Feedforward Networks
Peter Belcak, Roger Wattenhofer
We break the linear link between the layer size and its inference cost by introducing the fast feedforward (FFF) architecture, a log-time alternative to feedforward networks. We de…
Neural Combinatorial Logic Circuit Synthesis from Input-Output Examples
Peter Belcak, Roger Wattenhofer
We propose a novel, fully explainable neural approach to synthesis of combinatorial logic circuits from input-output examples. The carrying advantage of our method is that it readi…
A Neural Model for Regular Grammar Induction
Peter Belcák, David Hofer, Roger Wattenhofer
Grammatical inference is a classical problem in computational learning theory and a topic of wider influence in natural language processing. We treat grammars as a model of computa…