activity
20242026
collaborators

17 papers

cs.CL2026

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…

cs.CL2026

HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings

Rasmus Aavang, Rasmus T. Aavang, Giovanni Rizzi +5

Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Business Reporting Language (iXBRL) is mandat…

cs.CL2026

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

Mike Zhang, Ali Basirat, Desmond Elliott

Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in…

cs.CL2026

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls

Rasmus T. Aavang, Rasmus Tjalk-Bøggild, Alexandre Iolov +3

Earnings calls are a key source of financial information about public companies. However, extracting information from these calls is difficult. Unlike the templatic filings require…

cs.CL2026

Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality

Mike Zhang, Johannes Bjerva, Russa Biswas

We introduce fs1, a simple yet effective method that improves the factuality of reasoning traces by collecting them from large reasoning models and grounding them in knowledge grap…

cs.CL2026

WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain

Matthias De Lange, Warre Veys, Federico Retyk +16

Today's evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language processing (NLP) capabilities…