1 citations · 2 across the 8 of their papers we have counts for
4 papers · 1 filter
AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP
Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai +17
Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, ge…
How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
Longju Bai, Zhemin Huang, Xingyao Wang +5
The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of t…
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
Ruiqi He, Yushu He, Longju Bai +7
Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address thi…
Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba
Ruiqi He, Yushu He, Longju Bai +6
Existing humor datasets and evaluations predominantly focus on English, lacking resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, w…