Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
Xinwang Chen, Ning Liu, Yichen Zhu +2
Transformer-based Diffusion Probabilistic Models (DPMs) have shown more potential than CNN-based DPMs, yet their extensive computational requirements hinder widespread practical ap…
cs.CV2024
Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
Minjie Zhu, Yichen Zhu, Xin Liu +7
Multimodal Large Language Models (MLLMs) have showcased impressive skills in tasks related to visual understanding and reasoning. Yet, their widespread application faces obstacles…
cs.CV2024
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Yichen Zhu, Minjie Zhu, Ning Liu +3
In this paper, we introduce LLaVA- (LLaVA-Phi), an efficient multi-modal assistant that harnesses the power of the recently advanced small language model, Phi-2, to facilitate…