papers

Publications (10)

cs.CR2023

Unique Identification of 50,000+ Virtual Reality Users from Head & Hand Motion Data

Vivek Nair, Wenbo Guo, Justus Mattern +4

With the recent explosive growth of interest and investment in virtual reality (VR) and the so-called "metaverse," public attention has rightly shifted toward the unique security a…

cs.LG2025

INTELLECT-3: Technical Report

Prime Intellect Team, Mika Senghaas, Fares Obeid +20

We present INTELLECT-3, a 106B-parameter Mixture-of-Experts model (12B active) trained with large-scale reinforcement learning on our end-to-end RL infrastructure stack. INTELLECT-…

cs.CL2023

Psychologically-Inspired Causal Prompts

Zhiheng Lyu, Zhijing Jin, Justus Mattern +3

NLP datasets are richer than just input-output pairs; rather, they carry causal relations between the input and output variables. In this work, we take sentiment classification as…

cs.CL2022

Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement Patterns

Daniel Wiechmann, Yu Qiao, Elma Kerz +1

There is a growing interest in the combined use of NLP and machine learning methods to predict gaze patterns during naturalistic reading. While promising results have been obtained…

cs.LG2022

Differentially Private Language Models for Secure Data Sharing

Justus Mattern, Zhijing Jin, Benjamin Weggenmann +2

To protect the privacy of individuals whose data is being shared, it is of high importance to develop methods allowing researchers and companies to release textual data while provi…

cs.LG2025

INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning

Prime Intellect Team, Sami Jaghouar, Justus Mattern +11

We introduce INTELLECT-2, the first globally distributed reinforcement learning (RL) training run of a 32 billion parameter language model. Unlike traditional centralized training…

cs.CL2024

Smaller Language Models are Better Black-box Machine-Generated Text Detectors

Niloofar Mireshghallah, Justus Mattern, Sicun Gao +2

With the advent of fluent generative language models that can produce convincing utterances very similar to those written by humans, distinguishing whether a piece of text is machi…

cs.CL2023

Membership Inference Attacks against Language Models via Neighbourhood Comparison

Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin +3

Membership Inference attacks (MIAs) aim to predict whether a data sample was present in the training data of a machine learning model or not, and are widely used for assessing the…

cs.CR2022

The Limits of Word Level Differential Privacy

Justus Mattern, Benjamin Weggenmann, Florian Kerschbaum

As the issues of privacy and trust are receiving increasing attention within the research community, various attempts have been made to anonymize textual data. A significant subset…

cs.CL2025

Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias

Yuen Chen, Vethavikashini Chithrra Raghuram, Justus Mattern +2

Generated texts from large language models (LLMs) have been shown to exhibit a variety of harmful, human-like biases against various demographics. These findings motivate research…