1 citations · 1 across the 4 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
Ameen Ali, Shahar Katz, Lior Wolf +1
Large language models (LLMs) often develop learned mechanisms specialized to specific datasets, such as reliance on domain-specific correlations, which yield high-confidence predic…
cs.CL2024★ 1 cited
Segment-Based Attention Masking for GPTs
Shahar Katz, Liran Ringel, Yaniv Romano +1
Modern Language Models (LMs) owe much of their success to masked causal attention, the backbone of Generative Pre-Trained Transformer (GPT) models. Although GPTs can process the en…
cs.CL2024
Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
Shahar Katz, Lior Wolf
The success of Transformer-based Language Models (LMs) stems from their attention mechanism. While this mechanism has been extensively studied in explainability research, particula…