From the 1 of 4 linked papers with an AI index.
4 papers
Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages
Hielke Muizelaar, Giulia Rivetti, Marco Spruit +1
The paper investigates using machine translation to convert low‑resource language data into English and then fine‑tuning an English BERT model, comparing this approach to native‑la…
OpenExtract: Automated Data Extraction for Systematic Reviews in Health
Jim Achterberg, Bram Van Dijk, Jing Meng +8
This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLM…
XGenBoost: Synthesizing Small and Large Tabular Datasets with XGBoost
Jim Achterberg, Marcel Haas, Bram van Dijk +1
Tree ensembles such as XGBoost are often preferred for discriminative tasks in mixed-type tabular data, due to their inductive biases, minimal hyperparameter tuning, and training e…
The Data Sharing Paradox of Synthetic Data in Healthcare
Jim Achterberg, Bram van Dijk, Saif ul Islam +6
Synthetic data offers a promising solution to privacy concerns in healthcare by generating useful datasets in a privacy-aware manner. However, although synthetic data is typically…