3 papers
cs.LG2026
Data-Aware Random Feature Kernel for Transformers
Amirhossein Farzam, Hossein Mobahi, Nolan Andrew Miller +1
Transformers excel across domains, yet their quadratic attention complexity poses a barrier to scaling. Random-feature attention, as in Performers, can reduce this cost to linear i…
cs.AI2026
Hiding in Plain Text: Detecting Concealed Jailbreaks via Activation Disentanglement
Amirhossein Farzam, Majid Behabahani, Mani Malek +2
Large language models (LLMs) remain vulnerable to jailbreak prompts that are fluent and semantically coherent, and therefore difficult to detect with standard heuristics. A particu…
cs.CV2025
SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers
Oded Schlesinger, Amirhossein Farzam, J. Matias Di Martino +1
While Vision Transformers (ViT) have demonstrated remarkable performance across diverse tasks, their computational demands are substantial, scaling quadratically with the number of…