4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.CL2024★ 4 cited
Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
Jamba Team, Barak Lenz, Alan Arazi +58
We present Jamba-1.5, new instruction-tuned large language models based on our Jamba architecture. Jamba is a hybrid Transformer-Mamba mixture of experts architecture, providing hi…
cs.CL2022
Data Contamination: From Memorization to Exploitation
Inbal Magar, Roy Schwartz
Pretrained language models are typically trained on massive web-based datasets, which are often "contaminated" with downstream test sets. It is not clear to what extent models expl…