1 paper
So Hasegawa, Shailaja Keyur Sampat, Lei Liu +1
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables…