5 papers
Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering
Jingchen Sun, Shaobo Han, Ruiyi Zhang +5
Multimodal retrievers are essential for knowledge-based visual question answering, where they retrieve external evidence for image-question pairs. However, existing contrastive tra…
Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
Jingchen Sun, Shaobo Han, Deep Patel +3
Knowledge distillation establishes a learning paradigm that leverages both data supervision and teacher guidance. However, determining the optimal balance between learning from dat…
KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
Rohan Sharma, Dancheng Liu, Jingchen Sun +4
With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to persona…
CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
Jingchen Sun, Shaobo Han, Wataru Kohno +1
Contrastive Language-Audio Pretraining (CLAP) models have demonstrated unprecedented performance in various acoustic signal recognition tasks. Fiber-optic-based acoustic recognitio…
Craft: Cross-modal Aligned Features Improve Robustness of Prompt Tuning
Jingchen Sun, Rohan Sharma, Vishnu Suresh Lokhande +1
Prompt Tuning has emerged as a prominent research paradigm for adapting vision-language models to various downstream tasks. However, recent research indicates that prompt tuning me…