4 papers · 1 filter
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Xiang An, Yin Xie, Feilong Tang +27
We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of mu…
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
Bo Li, Ronghao Chen, Ningyuan Deng +3
Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within social media and e-commerce doma…
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
Sa Zhu, Wanqian Zhang, Lin Wang +3
Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effective alignment between action…
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation
Bo Li, Shaolin Zhu, Lijie Wen
Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages. However, existing datasets…