Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
DistillLens: Symmetric Knowledge Distillation Through Logit Lens
Manish Dhakal, Uthman Jinadu, Anjila Budathoki +2
Standard Knowledge Distillation (KD) compresses Large Language Models (LLMs) by optimizing final outputs, yet it typically treats the teacher's intermediate layer's thought process…
cs.CL2026
Can a Unimodal Language Agent Provide Preferences to Tune a Multimodal Vision-Language Model?
Sazia Tabasum Mim, Jack Morris, Manish Dhakal +3
To explore a more scalable path for adding multimodal capabilities to existing LLMs, this paper addresses a fundamental question: Can a unimodal LLM, relying solely on text, reason…