collaborators

6 papers

cs.CL2026

ToxiREX: A Dataset on Toxic REasoning in ConteXt

Stefan F. Schouten, Ilia Markov, Piek Vossen

We introduce a new, contextual, multilingual dataset called ToxiREX: Toxic REasoning in ConteXt. The dataset consists of threads of Reddit comments and structured characterizations…

cs.CL2025

Truth-value judgment in language models: 'truth directions' are context sensitive

Stefan F. Schouten, Peter Bloem, Ilia Markov +1

Recent work has demonstrated that the latent spaces of large language models (LLMs) contain directions predictive of the truth of sentences. Multiple methods recover such direction…

cs.LG2025

Layer-wise Quantization for Quantized Optimistic Dual Averaging

Anh Duc Nguyen, Ilia Markov, Frank Zhengqing Wu +4

Modern deep neural networks exhibit heterogeneity across numerous layers of various types such as residuals, multi-head attention, etc., due to varying structures (dimensions, acti…

cs.LG2025

Wasserstein Distances, Neuronal Entanglement, and Sparsity

Shashata Sawmya, Linghao Kong, Ilia Markov +2

Disentangling polysemantic neurons is at the core of many current approaches to interpretability of large language models. Here we attempt to study how disentanglement can be used…

cs.CL2025

Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant

Armand Nicolicioiu, Eugenia Iofinova, Andrej Jovanovic +6

The availability of powerful open-source large language models (LLMs) opens exciting use-cases, such as using personal data to fine-tune these models to imitate a user's unique wri…

cs.CL2025

Leveraging Open-Source Large Language Models for Native Language Identification

Yee Man Ng, Ilia Markov

Native Language Identification (NLI) - the task of identifying the native language (L1) of a person based on their writing in the second language (L2) - has applications in forensi…