6 papers
How to Correctly Report LLM-as-a-Judge Evaluations
Chungpa Lee, Thomas Zeng, Jongwon Jeong +2
Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensitivity and specificity of the LLM judges…
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
Dongmin Park, Minkyu Kim, Beongjun Choi +13
Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practica…
LLM-Lasso: A Robust Framework for Domain-Informed Feature Selection and Regularization
Erica Zhang, Ryunosuke Goto, Naomi Sagan +7
We introduce LLM-Lasso, a novel framework that leverages large language models (LLMs) to guide feature selection in Lasso regression. Unlike traditional methods that rely…
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
Anshumann, Mohd Abbas Zaidi, Akhil Kedia +5
Knowledge distillation can be a cost-effective technique to distill knowledge in Large Language Models, if the teacher output logits can be pre-computed and cached. However, succes…
Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents
Soyeon Choi, Kangwook Lee, Oliver Sng +1
How does the threat of infectious disease influence sociality among generative agents? We used generative agent-based modeling (GABM), powered by large language models, to experime…
Forward and Inverse Simulation of Pseudo-Two-Dimensional Model of Lithium-Ion Batteries Using Neural Networks
Myeong-Su Lee, Jaemin Oh, Dong-Chan Lee +3
In this work, we address the challenges posed by the high nonlinearity of the Butler-Volmer (BV) equation in forward and inverse simulations of the pseudo-two-dimensional (P2D) mod…