activity
20242026
most citedLessons from the Trenches on Reproducible Evaluation of Language Models

5 citations · 5 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL20265 cited

Lessons from the Trenches on Reproducible Evaluation of Language Models

Stella Biderman, Hailey Schoelkopf, Lintang Sutawika +27

Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivity of models to evaluation setup…

cs.CL2025

Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving

Xiangru Tang, Tianrui Qin, Tianhao Peng +15

AI agent frameworks operate in isolation, forcing agents to rediscover solutions and repeat mistakes across different systems. Despite valuable problem-solving experiences accumula…

cs.CL2025

ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning

Xiangru Tang, Tianyu Hu, Muyang Ye +9

Chemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large langu…

cs.CL2024

FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents

Yilun Zhao, Yitao Long, Yuru Jiang +7

We introduce FinDVer, a comprehensive benchmark specifically designed to evaluate the explainable claim verification capabilities of LLMs in the context of understanding and analyz…

cs.CL2024

ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Xiangru Tang, Yuliang Liu, Zefan Cai +21

Despite Large Language Models (LLMs) like GPT-4 achieving impressive results in function-level code generation, they struggle with repository-scale code understanding (e.g., coming…

cs.CL2024

Step-Back Profiling: Distilling User History for Personalized Scientific Writing

Xiangru Tang, Xingyao Zhang, Yanjun Shao +6

Large language models (LLM) excel at a variety of natural language processing tasks, yet they struggle to generate personalized content for individuals, particularly in real-world…