2 papers
cs.CV2025
Valley: Video Assistant with Large Language model Enhanced abilitY
Ruipu Luo, Ziwang Zhao, Min Yang +6
Large Language Models (LLMs), with remarkable conversational capability, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiven…
cs.CV2024
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
Zhen Wang, Da Li, Yulin Su +3
Logo embedding models convert the product logos in images into vectors, enabling their utilization for logo recognition and detection within e-commerce platforms. This facilitates…