2 papers
cs.CL2024
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
Chuyu Zhang, Songyang Zhang, Yingfan Hu +8
While LLM-Based agents, which use external tools to solve complex problems, have made significant progress, benchmarking their ability is challenging, thereby hindering a clear und…
cs.AI2024
Minor DPO reject penalty to increase training robustness
Shiming Xie, Hong Chen, Fred Yu +3
Learning from human preference is a paradigm used in large-scale language model (LLM) fine-tuning step to better align pretrained LLM to human preference for downstream task. In th…