2 papers
cs.IR2026
Unified Multimodal and Multilingual Retrieval via Multi-Task Learning with NLU Integration
Xinyuan Zhang, Lina Zhang, Lisung Chen +6
Multimodal retrieval systems typically employ Vision Language Models (VLMs) that encode images and text independently into vectors within a shared embedding space. Despite incorpor…
cs.CV2025
OregairuChar: A Benchmark Dataset for Character Appearance Frequency Analysis in My Teen Romantic Comedy SNAFU
Qi Sun, Dingju Zhou, Lina Zhang
The analysis of character appearance frequency is essential for understanding narrative structure, character prominence, and story progression in anime. In this work, we introduce…