collaborators

8 papers

cs.CV2026

The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh +2

Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal…

cs.LG2026

WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention

Ruben Solozabal, Velibor Bojkovic, Hilal Alquabeh +3

State-space models (SSMs) have emerged as a powerful foundation for long-range sequence modeling, with the HiPPO framework showing that continuous-time projection operators can be…

cs.CL2026

Sycophancy Hides Linearly in the Attention Heads

Rifo Genadi, Munachiso Nwadike, Nurdaulet Mukhituly +3

We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we trai…

cs.LG2025

Uncovering the Spectral Bias in Diagonal State Space Models

Ruben Solozabal, Velibor Bojkovic, Hilal AlQuabeh +2

Current methods for initializing state space models (SSMs) parameters mainly rely on the \textit{HiPPO framework}, which is based on an online approximation of orthogonal polynomia…

cs.CL2025

Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View

Muhammad Cendekia Airlangga, Hilal AlQuabeh, Munachiso S Nwadike +1

We study memory in state-space language models using primacy and recency effects as behavioral tools to uncover how information is retained and forgotten over time. Applying struct…

cs.LG2025

Mechanistic Insights into Grokking from the Embedding Layer

H. V. AlquBoj, Hilal AlQuabeh, Velibor Bojkovic +2

Grokking, a delayed generalization in neural networks after perfect training performance, has been observed in Transformers and MLPs, but the components driving it remain underexpl…