4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.LG2024
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
Roi Cohen, Konstantin Dobler, Eden Biran +1
Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are…
cs.CL2024★ 4 cited
Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
Jamba Team, Barak Lenz, Alan Arazi +58
We present Jamba-1.5, new instruction-tuned large language models based on our Jamba architecture. Jamba is a hybrid Transformer-Mamba mixture of experts architecture, providing hi…