papers
Publications (3)
cs.LG2026
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
cs.CL2022
ALIGN-MLM: Word Embedding Alignment is Crucial for Multilingual Pre-training
Henry Tang, Ameet Deshpande, Karthik Narasimhan
Multilingual pre-trained models exhibit zero-shot cross-lingual transfer, where a model fine-tuned on a source language achieves surprisingly good performance on a target language.…
cs.SE2021
On Using Stack Overflow Comment-Edit Pairs to recommend code maintenance changes
Henry Tang, Sarah Nadi
Code maintenance data sets typically consist of a before and after version of the code that contains the improvement or fix. Such data sets are important for software engineering s…