13 papers
T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning
Yanjun Fu, Faisal Hamman, Sanghamitra Dutta
Instruction tuning is essential for Large Language Models (LLMs) to effectively follow user instructions. To improve training efficiency and reduce data redundancy, recent works us…
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
Pasan Dissanayake, Sanghamitra Dutta
Transformer-based models have shown promising performance on tabular data compared to their classical counterparts such as neural networks and Gradient Boosted Decision Trees (GBDT…
Towards Formalizing Spuriousness of Biased Datasets Using Partial Information Decomposition
Barproda Halder, Faisal Hamman, Pasan Dissanayake +3
Spuriousness arises when there is an association between two or more variables in a dataset that are not causally related. In this work, we propose an explainability framework to p…
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
Faisal Hamman, Pasan Dissanayake, Yanjun Fu +1
Knowledge distillation is a promising approach to transfer capabilities from complex teacher models to smaller, resource-efficient student models that can be deployed easily, parti…
Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
Faisal Hamman, Chenyang Zhu, Anoop Kumar +4
RAG systems are increasingly deployed in high-stakes domains where users expect outputs to be consistent across semantically equivalent queries. However, existing systems often exh…
What If, But Privately: Private Counterfactual Retrieval
Shreya Meel, Mohamed Nomeir, Pasan Dissanayake +2
Transparency and explainability are two important aspects to be considered when employing black-box machine learning models in high-stake applications. Providing counterfactual exp…