works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators

14 papers

cs.MM2026

Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

Shiyu Li, Zhiyuan Hu, Yifan Wang +3

The paper presents Conan-embedding-v3, a framework that trains modality‑specific specialist models, fuses them into a single dense backbone, and then recovers the modality projecto…

cs.CV2026

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space

Peiming Li, Yifan Wang, Xiaotian Zhang +4

Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduct…

cs.CV2026

Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding

Yifan Wang, Peiming Li, Shiyu Li +5

While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…

cs.IR2026

TriAlignGR: Triangular Multitask Alignment with Multimodal Deep Interest Mining for Generative Recommendation

Yangchen Zeng, Hao Peng, Rongfeng Guo +3

We introduce TriAlignGR, a unified multitask-multimodal framework for generative recommendation that establishes two-stage multimodal semantic propagation: (i) encoding visual sema…

cs.CV2026

READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling

Thong Nguyen, Xiaobao Wu, Xinshuai Dong +5

Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal language grounding and video-language…

cs.CV2026

Multi-Scale Contrastive Learning for Video Temporal Grounding

Thong Thanh Nguyen, Yi Bin, Xiaobao Wu +4

Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video understanding. To encode video moment…