6 papers
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
Xiaochuan Li, Ke Wang, Girija Gouda +5
As Large Language Models (LLMs) become integrated into high-stakes domains, there is a growing need for evaluation methods that are both scalable for real-time deployment and relia…
Interpretable LLM-based Table Question Answering
Giang Nguyen, Ivan Brugere, Shubham Sharma +3
Interpretability in Table Question Answering (Table QA) is critical, especially in high-stakes domains like finance and healthcare. While recent Table QA approaches based on Large…
Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs
Faisal Hamman, Pasan Dissanayake, Saumitra Mishra +2
Fine-tuning LLMs on tabular classification tasks can lead to the phenomenon of fine-tuning multiplicity where equally well-performing models make conflicting predictions on the sam…
Interpreting Language Reward Models via Contrastive Explanations
Junqi Jiang, Tom Bewley, Saumitra Mishra +2
Reward models (RMs) are a crucial component in the alignment of large language models' (LLMs) outputs with human values. RMs approximate human preferences over possible LLM respons…
Graphusion: A RAG Framework for Knowledge Graph Construction with a Global Perspective
Rui Yang, Boming Yang, Aosong Feng +7
Knowledge Graphs (KGs) are crucial in the field of artificial intelligence and are widely used in downstream tasks, such as question-answering (QA). The construction of KGs typical…
Sequential Harmful Shift Detection Without Labels
Salim I. Amoukou, Tom Bewley, Saumitra Mishra +3
We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requi…