activity
20232026
most citedThe Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

6 citations · 24 across the 36 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Rethinking the Harmonic Loss via Non-Euclidean Distance Layers

Maxwell Miller-Golub, Collin Coil, Kamil Faber +4

Cross-entropy loss has long been the standard choice for training deep neural networks, yet it suffers from interpretability limitations, unbounded weight growth, and inefficiencie…

cs.LG2025

Universal Properties of Activation Sparsity in Modern Large Language Models

Filip Szatkowski, Patryk Będkowski, Alessio Devoto +5

Activation sparsity is an intriguing property of deep neural networks that has been extensively studied in ReLU-based models, due to its advantages for efficiency, robustness, and…

cs.LG2025

Neurosymbolic Reasoning Shortcuts under the Independence Assumption

Emile van Krieken, Pasquale Minervini, Edoardo Ponti +1

The ubiquitous independence assumption among symbolic concepts in neurosymbolic (NeSy) predictors is a convenient simplification: NeSy predictors use it to speed up probabilistic r…

cs.LG2025

Neurosymbolic Diffusion Models

Emile van Krieken, Pasquale Minervini, Edoardo Ponti +1

Neurosymbolic (NeSy) predictors combine neural perception with symbolic reasoning to solve tasks like visual reasoning. However, standard NeSy predictors assume conditional indepen…

cs.LG2024

When Can Proxies Improve the Sample Complexity of Preference Learning?

Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi +4

We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as…

cs.LG2024

An Auditing Test To Detect Behavioral Shift in Language Models

Leo Richter, Xuanli He, Pasquale Minervini +1

As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task perf…