Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding
Kaidong Feng, Zhuoxuan Huang, Huizhong Guo +7
Fashion understanding requires both visual perception and expert-level reasoning about style, occasion, compatibility, and outfit rationale. However, existing fashion datasets rema…
cs.CV2023
LMEye: An Interactive Perception Network for Large Language Models
Yunxin Li, Baotian Hu, Xinyu Chen +3
Training a Multimodal Large Language Model (MLLM) from scratch, like GPT-4, is resource-intensive. Regarding Large Language Models (LLMs) as the core processor for multimodal infor…