papers

Publications (56)

cs.LG2022

Set Interdependence Transformer: Set-to-Sequence Neural Networks for Permutation Learning and Structure Prediction

Mateusz Jurewicz, Leon Derczynski

The task of learning to map an input set onto a permuted sequence of its elements is challenging for neural networks. Set-to-sequence problems occur in natural language processing,…

cs.CL2025

Llama-Nemotron: Efficient Reasoning Models

Akhiad Bercovich, Itay Levy, Izik Golan +132

We introduce the Llama-Nemotron series of models, an open family of heterogeneous reasoning models that deliver exceptional reasoning capabilities, inference efficiency, and an ope…

cs.CL2024

garak: A Framework for Security Probing Large Language Models

Leon Derczynski, Erick Galinkin, Jeffrey Martin +2

As Large Language Models (LLMs) are deployed and integrated into thousands of applications, the need for scalable evaluation of how models respond to adversarial attacks grows rapi…

cs.CL2023

Assessing Language Model Deployment with Risk Cards

Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran +4

This paper introduces RiskCards, a framework for structured assessment and documentation of risks associated with an application of language models. As with all language, text gene…

cs.CL2017

SemEval-2017 Task 8: RumourEval: Determining rumour veracity and support for rumours

Leon Derczynski, Kalina Bontcheva, Maria Liakata +3

Media is full of false claims. Even Oxford Dictionaries named "post-truth" as the word of 2016. This makes it more important than ever to build systems that can identify the veraci…

cs.CL2023

Offensive Language and Hate Speech Detection for Danish

Gudbjartur Ingi Sigurbergsson, Leon Derczynski

The presence of offensive language on social media platforms and the implications this poses is becoming a major concern in modern society. Given the enormous amount of content cre…

cs.CL2016

Desiderata for Vector-Space Word Representations

Leon Derczynski

A plethora of vector-space representations for words is currently available, which is growing. These consist of fixed-length vectors containing real values, which represent a word.…

cs.CL2023

Discriminating Between Similar Nordic Languages

René Haas, Leon Derczynski

Automatic language identification is a challenging problem. Discriminating between closely related languages is especially difficult. This paper presents a machine learning approac…

cs.CL2019

Simple Natural Language Processing Tools for Danish

Leon Derczynski

This technical note describes a set of baseline tools for automatic processing of Danish text. The tools are machine-learning based, using natural language processing models traine…

cs.CL2012

A Data Driven Approach to Query Expansion in Question Answering

Leon Derczynski, Jun Wang, Robert Gaizauskas +1

Automated answering of natural language questions is an interesting and useful problem to solve. Question answering (QA) systems often perform information retrieval at an initial s…

cs.CL2012

An Annotation Scheme for Reichenbach's Verbal Tense Structure

Leon Derczynski, Robert Gaizauskas

In this paper we present RTMML, a markup language for the tenses of verbs and temporal relations between verbs. There is a richness to tense in language that is not fully captured…

cs.CL2012

Using Signals to Improve Automatic Classification of Temporal Relations

Leon Derczynski, Robert Gaizauskas

Temporal information conveyed by language describes how the world around us changes through time. Events, durations and times are all temporal elements that can be viewed as interv…

cs.CL2015

USFD: Twitter NER with Drift Compensation and Linked Data

Leon Derczynski, Isabelle Augenstein, Kalina Bontcheva

This paper describes a pilot NER system for Twitter, comprising the USFD system entry to the W-NUT 2015 NER shared task. The goal is to correctly label entities in a tweet dataset,…

cs.CL2025

NLP Security and Ethics, in the Wild

Heather Lent, Erick Galinkin, Yiyi Chen +3

As NLP models are used by a growing number of end-users, an area of increasing importance is NLP Security (NLPSec): assessing the vulnerability of models to malicious attacks and d…

cs.LG2026

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

NVIDIA, :, Amala Sanjay Deshmukh +204

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 N…

cs.CR2026

Training a General Purpose Automated Red Teaming Model

Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1

Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They…

cs.LG2025

Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities

Arjun Krishna, Erick Galinkin, Leon Derczynski +1

Large Language Models (LLMs) have become an essential tool in the programmer's toolkit, but their tendency to hallucinate code can be used by malicious actors to introduce vulnerab…

cs.CL2024

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97

This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…

cs.CL2013

TimeML-strict: clarifying temporal annotation

Leon Derczynski, Hector Llorens, Naushad UzZaman

TimeML is an XML-based schema for annotating temporal information over discourse. The standard has been used to annotate a variety of resources and is followed by a number of tools…

cs.CL2013

Question Answering Against Very-Large Text Collections

Leon Derczynski, Richard Shaw, Ben Solway +1

Question answering involves developing methods to extract useful information from large collections of documents. This is done with specialised search engines such as Answer Finder…

cs.CL2025

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

NVIDIA, :, Aaron Blakeman +198

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemo…

cs.CL2012

A Corpus-based Study of Temporal Signals

Leon Derczynski, Robert Gaizauskas

Automatic temporal ordering of events described in discourse has been of great interest in recent years. Event orderings are conveyed in text via va rious linguistic mechanisms inc…

cs.CL2024

Nemotron-4 340B Technical Report

Nvidia, :, Bo Adler +80

We release the Nemotron-4 340B model family, including Nemotron-4-340B-Base, Nemotron-4-340B-Instruct, and Nemotron-4-340B-Reward. Our models are open access under the NVIDIA Open…

cs.CL2023

Surveying (Dis)Parities and Concerns of Compute Hungry NLP Research

Ji-Ung Lee, Haritz Puerto, Betty van Aken +8

Many recent improvements in NLP stem from the development and use of large pre-trained language models (PLMs) with billions of parameters. Large model sizes makes computational cos…

cs.CL2012

USFD at KBP 2011: Entity Linking, Slot Filling and Temporal Bounding

Amev Burman, Arun Jayapal, Sathish Kannan +4

This paper describes the University of Sheffield's entry in the 2011 TAC KBP entity linking and slot filling tasks. We chose to participate in the monolingual entity linking task,…

cs.CL2025

NVIDIA Nemotron 3: Efficient and Open Intelligence

NVIDIA, :, Aaron Blakeman +356

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a…

cs.CL2026

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +571

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…

cs.CL2014

TempEval-3: Evaluating Events, Time Expressions, and Temporal Relations

Naushad UzZaman, Hector Llorens, James Allen +3

We describe the TempEval-3 task which is currently in preparation for the SemEval-2013 evaluation exercise. The aim of TempEval is to advance research on temporal information proce…

cs.CL2012

Massively Increasing TIMEX3 Resources: A Transduction Approach

Leon Derczynski, Héctor Llorens, Estela Saquete

Automatic annotation of temporal expressions is a research challenge of great interest in the field of information extraction. Gold standard temporally-annotated resources are limi…

cs.CL2023

Efficient Methods for Natural Language Processing: A Survey

Marcos Treviso, Ji-Ung Lee, Tianchu Ji +19

Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance mea…

cs.CL2014

Analysis of Named Entity Recognition and Linking for Tweets

Leon Derczynski, Diana Maynard, Giuseppe Rizzo +5

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully author…

cs.CL2023

Sparse Probability of Agreement

Jeppe Nørregaard, Leon Derczynski

Measuring inter-annotator agreement is important for annotation tasks, but many metrics require a fully-annotated set of data, where all annotators annotate all samples. We define…

cs.CL2018

Helping Crisis Responders Find the Informative Needle in the Tweet Haystack

Leon Derczynski, Kenny Meesters, Kalina Bontcheva +1

Crisis responders are increasingly using social media, data and other digital sources of information to build a situational understanding of a crisis situation in order to design a…

cs.CL2022

Bridging the Domain Gap for Stance Detection for the Zulu language

Gcinizwe Dlamini, Imad Eddine Ibrahim Bekkouch, Adil Khan +1

Misinformation has become a major concern in recent last years given its spread across our information sources. In the past years, many NLP tasks have been introduced in this area,…

cs.CL2020

SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)

Marcos Zampieri, Preslav Nakov, Sara Rosenthal +6

We present the results and main findings of SemEval-2020 Task 12 on Multilingual Offensive Language Identification in Social Media (OffensEval 2020). The task involves three subtas…

cs.CL2017

Generalisation in Named Entity Recognition: A Quantitative Analysis

Isabelle Augenstein, Leon Derczynski, Kalina Bontcheva

Named Entity Recognition (NER) is a key NLP task, which is all the more challenging on Web and user-generated content with their diverse and continuously changing language. This pa…

cs.CL2022

Detecting Abusive Albanian

Erida Nurce, Jorgel Keci, Leon Derczynski

The ever growing usage of social media in the recent years has had a direct impact on the increased presence of hate speech and offensive speech in online platforms. Research on ef…

cs.CL2012

Analysing Temporally Annotated Corpora with CAVaT

Leon Derczynski, Robert Gaizauskas

We present CAVaT, a tool that performs Corpus Analysis and Validation for TimeML. CAVaT is an open source, modular checking utility for statistical analysis of features specific to…

cs.CL2021

Optimal Size-Performance Tradeoffs: Weighing PoS Tagger Models

Magnus Jacobsen, Mikkel H. Sørensen, Leon Derczynski

Improvement in machine learning-based NLP performance are often presented with bigger models and more complex code. This presents a trade-off: better scores come at the cost of lar…

cs.CL2017

Tracking the Diffusion of Named Entities

Leon Derczynski, Matthew Rowe

Existing studies of how information diffuses across social networks have thus far concentrated on analysing and recovering the spread of deterministic innovations such as URLs, has…

cs.CL2017

Simple Open Stance Classification for Rumour Analysis

Ahmet Aker, Leon Derczynski, Kalina Bontcheva

Stance classification determines the attitude, or stance, in a (typically short) text. The task has powerful applications, such as the detection of fake news or the automatic extra…

cs.CL2021

Directions in Abusive Language Training Data: Garbage In, Garbage Out

Bertie Vidgen, Leon Derczynski

Data-driven analysis and detection of abusive online content covers many different tasks, phenomena, contexts, and methodologies. This paper systematically reviews abusive language…

cs.CL2023

Handling and Presenting Harmful Text in NLP Research

Hannah Rose Kirk, Abeba Birhane, Bertie Vidgen +1

Text data can pose a risk of harm. However, the risks are not fully understood, and how to handle, present, and discuss harmful text in a safe way remains an unresolved issue in th…

cs.LG2020

Power Consumption Variation over Activation Functions

Leon Derczynski

The power that machine learning models consume when making predictions can be affected by a model's architecture. This paper presents various estimates of power consumption for a r…

cs.CL2012

USFD2: Annotating Temporal Expresions and TLINKs for TempEval-2

Leon Derczynski, Robert Gaizauskas

We describe the University of Sheffield system used in the TempEval-2 challenge, USFD2. The challenge requires the automatic identification of temporal entities and relations in te…

cs.CL2025

NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

NVIDIA, :, Aarti Basant +214

We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compar…

cs.CL2024

Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming

Nanna Inie, Jonathan Stray, Leon Derczynski

Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition o…

cs.CL2022

Training a T5 Using Lab-sized Resources

Manuel R. Ciosici, Leon Derczynski

Training large neural language models on large datasets is resource- and time-intensive. These requirements create a barrier to entry, where those with fewer resources cannot build…

cs.CL2022

The ITU Faroese Pairs Dataset

Leon Derczynski, Annika Solveig Hedegaard Isfeldt, Signhild Djurhuus

This article documents a dataset of sentence pairs between Faroese and Danish, produced at ITU Copenhagen. The data covers tranlsation from both source languages, and is intended f…

cs.CL2020

The Rumour Mill: Making the Spread of Misinformation Explicit and Tangible

Nanna Inie, Jeanette Falk Olesen, Leon Derczynski

Misinformation spread presents a technological and social threat to society. With the advance of AI-based language models, automatically generated texts have become difficult to id…

cs.LG2026

Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aakshita Chandiramani +544

We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemo…

cs.CL2014

Clinical TempEval

Steven Bethard, Leon Derczynski, James Pustejovsky +1

We describe the Clinical TempEval task which is currently in preparation for the SemEval-2015 evaluation exercise. This task involves identifying and describing events, times and t…

cs.AI2026

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Paul Furgale, Severin Klingler, James Nolan +12

Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic…

cs.CL2018

RumourEval 2019: Determining Rumour Veracity and Support for Rumours

Genevieve Gorrell, Kalina Bontcheva, Leon Derczynski +3

This is the proposal for RumourEval-2019, which will run in early 2019 as part of that year's SemEval event. Since the first RumourEval shared task in 2017, interest in automated c…

cs.CL2025

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +311

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 t…

cs.CL2018

Stance Prediction for Russian: Data and Analysis

Nikita Lozhnikov, Leon Derczynski, Manuel Mazzara

Stance detection is a critical component of rumour and fake news identification. It involves the extraction of the stance a particular author takes related to a given claim, both e…