3 papers
cs.LG2026
Validity-Calibrated Reasoning Distillation
Khouloud Saadi, Di Wang
Reasoning distillation aims to transfer multi-step reasoning capabilities from large language models to smaller, more efficient ones. While recent methods have shown promising gain…
cs.CL2026
What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View
Khouloud Saadi, Di Wang
Feature-based knowledge distillation aims to transfer intermediate representations from a teacher LLM model to a student. Existing approaches typically rely on direct feature match…
cs.LG2025
Dissecting Representation Misalignment in Contrastive Learning via Influence Function
Lijie Hu, Chenyang Ren, Huanyi Xie +5
Contrastive learning, commonly applied in large-scale multimodal models, often relies on data from diverse and often unreliable sources, which can include misaligned or mislabeled…