3 papers
cs.LG2025
Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
Saeed AlMarri, Mathieu Ravaut, Kristof Juhasz +3
Large Language Models (LLMs) have attracted significant attention for classification tasks, offering a flexible alternative to trusted classical machine learning models like LightG…
cs.CL2025
Interpreting LLMs as Credit Risk Classifiers: Do Their Feature Explanations Align with Classical ML?
Saeed AlMarri, Kristof Juhasz, Mathieu Ravaut +3
Large Language Models (LLMs) are increasingly explored as flexible alternatives to classical machine learning models for classification tasks through zero-shot prompting. However,…
cs.CL2025
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
Hailin Chen, Fangkai Jiao, Mathieu Ravaut +8
The rapid advancement of large language models (LLMs) demands robust, unbiased, and scalable evaluation methods. However, human annotations are costly to scale, model-based evaluat…