2 papers
cs.CV2025
Learning to Instruct for Visual Instruction Tuning
Zhihan Zhou, Feng Hong, Jiaan Luo +5
We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for V…
cs.LG2025
Long-tailed Recognition with Model Rebalancing
Jiaan Luo, Feng Hong, Qiang Hu +3
Long-tailed recognition is ubiquitous and challenging in deep learning and even in the downstream finetuning of foundation models, since the skew class distribution generally preve…