3 papers
cs.CL2026
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
Nischay Dhankhar, Dos Baha, Abulhair Saparov
Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge i…
cs.CL2024
H2O-Danube3 Technical Report
Pascal Pfeiffer, Philipp Singer, Yauhen Babakhin +3
We present H2O-Danube3, a series of small language models consisting of H2O-Danube3-4B, trained on 6T tokens and H2O-Danube3-500M, trained on 4T tokens. Our models are pre-trained…
cs.CL2024
H2O-Danube-1.8B Technical Report
Philipp Singer, Pascal Pfeiffer, Yauhen Babakhin +4
We present H2O-Danube, a series of small 1.8B language models consisting of H2O-Danube-1.8B, trained on 1T tokens, and the incremental improved H2O-Danube2-1.8B trained on an addit…