activity
20242026
most citedLabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs

5 citations · 6 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL20265 cited

LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs

Yujun Zhou, Jingdong Yang, Yue Huang +12

Artificial Intelligence (AI) is revolutionizing scientific research, yet its growing integration into laboratory environments presents critical safety challenges. Large language mo…

cs.CL2025

LLMs4All: A Review of Large Language Models Across Academic Disciplines

Yanfang Ye, Zheyuan Zhang, Tianyi Ma +26

Cutting-edge Artificial Intelligence (AI) techniques keep reshaping our view of the world. For example, Large Language Models (LLMs) based applications such as ChatGPT have shown t…

cs.CL2025

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

Anna Sokol, Elizabeth Daly, Michael Hind +4

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation method…

cs.CL2024

NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning

Zheyuan Zhang, Yiyang Li, Nhi Ha Lan Le +9

Diet plays a critical role in human health, yet tailoring dietary reasoning to individual health conditions remains a major challenge. Nutrition Question Answering (QA) has emerged…

cs.CL2024

Class-Aware Contrastive Optimization for Imbalanced Text Classification

Grigorii Khvatskii, Nuno Moniz, Khoa Doan +1

The unique characteristics of text data make classification tasks a complex problem. Advances in unsupervised and semi-supervised learning and autoencoder architectures addressed s…

cs.CL2024

Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Jiayi Ye, Yanbo Wang, Yue Huang +9

LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many…