activity
20242026
collaborators

5 papers

cs.CL2026

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

Amit LeVi, Raz Lapid, Rom Himelstein +3

Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering the target task. This can waste…

cs.CR2026

Jailbreak Attack Initializations as Extractors of Compliance Directions

Amit Levi, Rom Himelstein, Yaniv Nemcovsky +2

Safety-aligned LLMs respond to prompts with either compliance or refusal, each corresponding to distinct directions in the model's activation space. Recent works show that initiali…

cs.CL2026

Silenced Biases: The Dark Side LLMs Learned to Refuse

Rom Himelstein, Amit LeVi, Brit Youngmann +2

Safety-aligned large language models (LLMs) are becoming increasingly widespread, especially in sensitive applications where fairness is essential and biased outputs can cause sign…

cs.CL2025

Representing LLMs in Prompt Semantic Task Space

Idan Kashani, Avi Mendelson, Yaniv Nemcovsky

Large language models (LLMs) achieve impressive results over various tasks, and ever-expanding public repositories contain an abundance of pre-trained models. Therefore, identifyin…

cs.CV2024

Sparse patches adversarial attacks via extrapolating point-wise information

Yaniv Nemcovsky, Avi Mendelson, Chaim Baskin

Sparse and patch adversarial attacks were previously shown to be applicable in realistic settings and are considered a security risk to autonomous systems. Sparse adversarial pertu…