6 papers
Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training
Yunbo Long, Tejumade Afonja, Guangya Hao +2
Tabular language models can generate synthetic tables by modeling rows as token sequences, but they are typically trained once with supervised fine-tuning and then used as static s…
The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes
Siobhan Mackenzie Hall, Samantha Dalal, Raesetje Sefala +8
This paper provides guidance for building and maintaining infrastructure for participatory AI efforts by sharing reflections on building World Wide Dishes (WWD), a bottom-up, commu…
Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches
Israel Abebe Azime, Deborah D. Kanubala, Tejumade Afonja +4
Large Language Models (LLMs) are increasingly employed in high-stakes decision-making tasks, such as loan approvals. While their applications expand across domains, LLMs struggle t…
The World Wide recipe: A community-centred framework for fine-grained data collection and regional bias operationalisation
Jabez Magomere, Shu Ishida, Tejumade Afonja +11
We introduce the World Wide recipe, which sets forth a framework for culturally aware and participatory data collection, and the resultant regionally diverse World Wide Dishes eval…
DP-2Stage: Adapting Language Models as Differentially Private Tabular Data Generators
Tejumade Afonja, Hui-Po Wang, Raouf Kerkouche +1
Generating tabular data under differential privacy (DP) protection ensures theoretical privacy guarantees but poses challenges for training machine learning models, primarily due t…
LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation
Tejumade Afonja, Ivaxi Sheth, Ruta Binkyte +4
Gene regulatory networks (GRNs) represent the causal relationships between transcription factors (TFs) and target genes in single-cell RNA sequencing (scRNA-seq) data. Understandin…