2 papers
cs.CV2024
VideoDistill: Language-aware Vision Distillation for Video Question Answering
Bo Zou, Chao Yang, Yu Qiao +2
Significant advancements in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although these image-language models c…
cs.CV2024
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
Bo Zou, Chao Yang, Yu Qiao +2
Existing methods to fine-tune LLMs, like Adapter, Prefix-tuning, and LoRA, which introduce extra modules or additional input sequences to inject new skills or knowledge, may compro…