5 citations · 6 across the 2 of their papers we have counts for
6 papers · 1 filter
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
Yujun Zhou, Jingdong Yang, Yue Huang +12
Artificial Intelligence (AI) is revolutionizing scientific research, yet its growing integration into laboratory environments presents critical safety challenges. Large language mo…
LLMs4All: A Review of Large Language Models Across Academic Disciplines
Yanfang Ye, Zheyuan Zhang, Tianyi Ma +26
Cutting-edge Artificial Intelligence (AI) techniques keep reshaping our view of the world. For example, Large Language Models (LLMs) based applications such as ChatGPT have shown t…
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
Anna Sokol, Elizabeth Daly, Michael Hind +4
Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation method…
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
Zheyuan Zhang, Yiyang Li, Nhi Ha Lan Le +9
Diet plays a critical role in human health, yet tailoring dietary reasoning to individual health conditions remains a major challenge. Nutrition Question Answering (QA) has emerged…
Class-Aware Contrastive Optimization for Imbalanced Text Classification
Grigorii Khvatskii, Nuno Moniz, Khoa Doan +1
The unique characteristics of text data make classification tasks a complex problem. Advances in unsupervised and semi-supervised learning and autoencoder architectures addressed s…
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Jiayi Ye, Yanbo Wang, Yue Huang +9
LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many…