activity
20242026
collaborators
Showing 2025Show all

8 papers · 1 filter

cs.LG2025

A Composable Channel-Adaptive Architecture for Seizure Classification

Francesco Carzaniga, Michael Hersche, Kaspar Schindler +1

Objective: We develop a channel-adaptive (CA) architecture that seamlessly processes multi-variate time-series with an arbitrary number of channels, and in particular intracranial…

cs.AI2025

Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models

Aleksandar Terzić, Nicolas Menet, Michael Hersche +2

Modern state-space models (SSMs) often utilize transition matrices which enable efficient computation but pose restrictions on the model's expressivity, as measured in terms of the…

cs.LG2025

Scalable Evaluation and Neural Models for Compositional Generalization

Giacomo Camposampiero, Pietro Barbiero, Michael Hersche +2

Compositional generalization-a key open challenge in modern machine learning-requires models to predict unknown combinations of known concepts. However, assessing compositional gen…

cs.LG2025

I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models

Giacomo Camposampiero, Michael Hersche, Roger Wattenhofer +2

We introduce I-RAVEN-X, a symbolic benchmark designed to evaluate generalization and robustness in analogical and mathematical reasoning for Large Language Models (LLMs) and Large…

cs.LG2025

A foundation model with multi-variate parallel attention to generate neuronal activity

Francesco Carzaniga, Michael Hersche, Abu Sebastian +2

Learning from multi-variate time-series with heterogeneous channel configurations remains a fundamental challenge for deep neural networks, particularly in clinical domains such as…

cs.LG2025

On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages

Aleksandar Terzić, Michael Hersche, Giacomo Camposampiero +3

Selective state-space models (SSMs) are an emerging alternative to the Transformer, offering the unique advantage of parallel training and sequential inference. Although these mode…