1 paper
Richard Zhuang, Akshat Gupta, Richard Yang +3
We introduce PokerBench - a benchmark for evaluating the poker-playing abilities of large language models (LLMs). As LLMs excel in traditional NLP tasks, their application to compl…