6 papers
The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage
Afnan Aloraini, Viktor Schlegel, Goran Nenadic +1
Automatic semantic change detection aims to identify how word meanings shift over time, offering insights into both linguistic and societal change. Despite recent progress in compu…
SynBench: A Benchmark for Differentially Private Text Generation
Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar +9
Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regu…
Evaluation and LLM-Guided Learning of ICD Coding Rationales
Mingyang Li, Viktor Schlegel, Tingting Mu +3
ICD coding is the process of mapping unstructured text from Electronic Health Records (EHRs) to standardised codes defined by the International Classification of Diseases (ICD) sys…
Term2Note: Synthesising Differentially Private Clinical Notes from Medical Terms
Yuping Wu, Viktor Schlegel, Warren Del-Pinto +10
Training data is fundamental to the success of modern machine learning models, yet in high-stakes domains such as healthcare, the use of real-world training data is severely constr…
Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs
Mingyang Li, Viktor Schlegel, Tingting Mu +2
Mapping clinical documents to standardised clinical vocabularies is an important task, as it provides structured data for information retrieval and analysis, which is essential to…
Evaluating Differentially Private Generation of Domain-Specific Text
Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar +7
Generative AI offers transformative potential for high-stakes domains such as healthcare and finance, yet privacy and regulatory barriers hinder the use of real-world data. To addr…