7 citations · 43 across the 33 of their papers we have counts for
7 papers · 2 filters
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls
Rasmus T. Aavang, Rasmus Tjalk-Bøggild, Alexandre Iolov +3
Earnings calls are a key source of financial information about public companies. However, extracting information from these calls is difficult. Unlike the templatic filings require…
CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
Mike Zhang, Ali Basirat, Desmond Elliott
Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in…
WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain
Matthias De Lange, Warre Veys, Federico Retyk +16
Today's evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language processing (NLP) capabilities…
UniSkill: A Dataset for Matching University Curricula to Professional Competencies
Nurlan Musazade, Joszef Mezei, Mike Zhang
Skill extraction and recommendation systems have been studied from recruiter, applicant, and education perspectives. While AI applications in job advertisements have received broad…
Do Large Language Models Adapt to Language Variation across Socioeconomic Status?
Elisa Bassignana, Mike Zhang, Dirk Hovy +1
Humans adjust their linguistic style to the audience they are addressing. However, the extent to which LLMs adapt to different social contexts is largely unknown. As these models i…