9 papers
SynBench: A Benchmark for Differentially Private Text Generation
Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar +9
Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regu…
Evaluation and LLM-Guided Learning of ICD Coding Rationales
Mingyang Li, Viktor Schlegel, Tingting Mu +3
ICD coding is the process of mapping unstructured text from Electronic Health Records (EHRs) to standardised codes defined by the International Classification of Diseases (ICD) sys…
MIRA: Medical Time Series Foundation Model for Real-World Health Data
Hao Li, Bowen Deng, Chang Xu +8
A unified foundation model for medical time series -- pretrained on open access and ethics board-approved medical corpora -- offers the potential to reduce annotation burdens, mini…
Large Language Models in Argument Mining: A Survey
Hao Li, Viktor Schlegel, Yizheng Sun +2
Large Language Models (LLMs) have fundamentally reshaped Argument Mining (AM), shifting it from a pipeline of supervised, task-specific classifiers to a spectrum of prompt-driven,…
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
Hao Li, Yizheng Sun, Viktor Schlegel +3
Argument summarization aims to generate concise, structured representations of complex, multi-perspective debates. While recent work has advanced the identification and clustering…
Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs
Mingyang Li, Viktor Schlegel, Tingting Mu +2
Mapping clinical documents to standardised clinical vocabularies is an important task, as it provides structured data for information retrieval and analysis, which is essential to…