1 paper
ZongHan Hsieh, Tzer-Jen Wei, ShengJing Yang
In this paper, we present ZonUI-3B, a lightweight Vision-Language Model (VLM) that can be fully trained on a single consumer-grade GPU (RTX 4090) while delivering performance compa…