learning efficiency 1multimodal alignment 1reasoning tasks 1text-to-vision mapping 1transformer fusion 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
Xuhui Zhan, Tyler Derr
The paper introduces Inverse-LLaVA, a multimodal model that projects text embeddings into continuous visual representation space and fuses them within transformer layers, reducing…
cs.IR2025
Towards Bridging Review Sparsity in Recommendation with Textual Edge Graph Representation
Leyao Wang, Xutao Mao, Xuhui Zhan +5
Textual reviews enrich recommender systems with fine-grained preference signals and enhanced explainability. However, in real-world scenarios, users rarely leave reviews, resulting…
cs.CV2025
Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment
Yang Hu, Runchen Wang, Stephen Chong Zhao +4
We introduce Perceptual-Initialization (PI), a paradigm shift in visual representation learning that incorporates human perceptual structure during the initialization phase rather…