From the 1 of 14 linked papers with an AI index.
14 papers
Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
Shiyu Li, Zhiyuan Hu, Yifan Wang +3
The paper presents Conan-embedding-v3, a framework that trains modality‑specific specialist models, fuses them into a single dense backbone, and then recovers the modality projecto…
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Peiming Li, Yifan Wang, Xiaotian Zhang +4
Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduct…
Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding
Yifan Wang, Peiming Li, Shiyu Li +5
While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…
TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation
Yangchen Zeng, Hao Peng, Rongfeng Guo +3
We introduce TriAlignGR, a unified multitask-multimodal framework for generative recommendation that establishes two-stage multimodal semantic propagation: (i) encoding visual sema…
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
Thong Nguyen, Xiaobao Wu, Xinshuai Dong +5
Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language…
Multi-Scale Contrastive Learning for Video Temporal Grounding
Thong Thanh Nguyen, Yi Bin, Xiaobao Wu +4
Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moment…