3 papers
cs.LG2026
Boosting Meta-Learning for Few-Shot Text Classification via Label-guided Distance Scaling
Yunlong Gao, Xinyue Liu, Yingbo Wang +2
Few-shot text classification aims to recognize unseen classes with limited labeled text samples. Existing approaches focus on boosting meta-learners by developing complex algorithm…
cs.CV2025
FILA: Fine-Grained Vision Language Models
Shiding Zhu, Wenhui Dong, Jun Song +3
Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common approach currently involves dyna…
cs.CV2025
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
Yanan Guo, Wenhui Dong, Jun Song +7
Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing…