26 citations · 92 across the 49 of their papers we have counts for
7 papers · 2 filters
BioLORD: Learning Ontological Representations from Definitions (for Biomedical Concepts and their Textual Descriptions)
François Remy, Kris Demuynck, Thomas Demeester
This work introduces BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts. State-of-the-art methodologies op…
Robustifying Sentiment Classification by Maximally Exploiting Few Counterfactuals
Maarten De Raedt, Fréderic Godin, Chris Develder +1
For text classification tasks, finetuned language models perform remarkably well. Yet, they tend to rely on spurious patterns in training data, thus limiting their performance on o…
EduQG: A Multi-format Multiple Choice Dataset for the Educational Domain
Amir Hadifar, Semere Kiros Bitew, Johannes Deleu +2
We introduce a high-quality dataset that contains 3,397 samples comprising (i) multiple choice questions, (ii) answers (including distractors), and (iii) their source documents, fr…
Learning to Reuse Distractors to support Multiple Choice Question Generation in Education
Semere Kiros Bitew, Amir Hadifar, Lucas Sterckx +3
Multiple choice questions (MCQs) are widely used in digital learning systems, as they allow for automating the assessment process. However, due to the increased digital literacy of…
Design of Negative Sampling Strategies for Distantly Supervised Skill Extraction
Jens-Joris Decorte, Jeroen Van Hautte, Johannes Deleu +2
Skills play a central role in the job market and many human resources (HR) processes. In the wake of other digital experiences, today's online job market has candidates expecting t…
Next-Year Bankruptcy Prediction from Textual Data: Benchmark and Baselines
Henri Arno, Klaas Mulier, Joke Baeck +1
Models for bankruptcy prediction are useful in several real-world scenarios, and multiple research contributions have been devoted to the task, based on structured (numerical) as w…