activity
20212026
most citedJamba-1.5: Hybrid Transformer-Mamba Models at Scale

4 citations · 4 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2026

A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models

Noa Linder, Meirav Segal, Omer Antverg +5

Large language models and LLM-based agents are increasingly used for cybersecurity tasks that are inherently dual-use. Existing approaches to refusal, spanning academic policy fram…

cs.CL2024★ 4 cited

Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Jamba Team, Barak Lenz, Alan Arazi +58

We present Jamba-1.5, new instruction-tuned large language models based on our Jamba architecture. Jamba is a hybrid Transformer-Mamba mixture of experts architecture, providing hi…

cs.CL2022

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…

cs.CL2022

IDANI: Inference-time Domain Adaptation via Neuron-level Interventions

Omer Antverg, Eyal Ben-David, Yonatan Belinkov

Large pre-trained models are usually fine-tuned on downstream task data, and tested on unseen data. When the train and test data come from different domains, the model is likely to…

cs.CL2021

On the Pitfalls of Analyzing Individual Neurons in Language Models

Omer Antverg, Yonatan Belinkov

While many studies have shown that linguistic information is encoded in hidden word representations, few have studied individual neurons, to show how and in which neurons it is enc…