activity
20242026
collaborators

10 papers

cs.LG2026

Excitation: Momentum For Experts

Sagi Shaier

We propose Excitation, a novel optimization framework designed to accelerate learning in sparse architectures such as Mixture-of-Experts (MoEs). Unlike traditional optimizers that…

cs.CL2025

MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset

Sagi Shaier, George Arthur Baker, Chiranthan Sridhar +2

Language models (LMs) have excelled in various broad domains. However, to ensure their safe and effective integration into real-world educational settings, they must demonstrate pr…

cs.CL2025

Asking Again and Again: Exploring LLM Robustness to Repeated Questions

Sagi Shaier, Mario Sanz-Guerrero, Katharina von der Wense

This study investigates whether repeating questions within prompts influences the performance of large language models (LLMs). We hypothesize that reiterating a question within a s…

cs.LG2025

More Experts Than Galaxies: Conditionally-overlapping Experts With Biologically-Inspired Fixed Routing

Sagi Shaier, Francisco Pereira, Katharina von der Wense +2

The evolution of biological neural systems has led to both modularity and sparse coding, which enables energy efficiency and robustness across the diversity of tasks in the lifespa…

cs.CL2024

Lost in the Middle, and In-Between: Enhancing Language Models' Ability to Reason Over Long Contexts in Multi-Hop QA

George Arthur Baker, Ankush Raut, Sagi Shaier +2

Previous work finds that recent long-context language models fail to make equal use of information in the middle of their inputs, preferring pieces of information located at the ta…

cs.CL2024

Comparing Template-based and Template-free Language Model Probing

Sagi Shaier, Kevin Bennett, Lawrence E Hunter +1

The differences between cloze-task language model (LM) probing with 1) expert-made templates and 2) naturally-occurring text have often been overlooked. Here, we evaluate 16 differ…