activity
20242026
collaborators

5 papers

cs.CL2026

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

Nithin Somasekharan, Ling Yue, Yadi Cao +6

Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experiments of complex physical system -- a cr…

cs.LG2025

Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models

Rihui Jin, Zheyu Xin, Xing Xie +6

Table reasoning (TR) requires structured reasoning over semi-structured tabular data and remains challenging, particularly for small language models (SLMs, e.g., LLaMA-8B) due to t…

cs.AI2025

General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…

cs.CL2025

Uncovering inequalities in new knowledge learning by large language models across different languages

Chenglong Wang, Haoyu Tang, Xiyuan Yang +8

As large language models (LLMs) gradually become integral tools for problem solving in daily life worldwide, understanding linguistic inequality is becoming increasingly important.…

cs.AI2024

Large Language Models show both individual and collective creativity comparable to humans

Luning Sun, Yuzhuo Yuan, Yuan Yao +6

Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to hum…