4 papers
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
Jen-Tse Huang, Dasen Dai, Jen-Yuan Huang +7
Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, current Multimodal Large Language Models…
Long-Text-to-Image Generation via Compositional Prompt Decomposition
Jen-Yuan Huang, Tong Lin, Yilun Du
While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are descriptive paragraphs. This…
Dynamic Pyramid Network for Efficient Multimodal Large Language Model
Hao Ai, Kunyi Wang, Zezhou Wang +7
Multimodal large language models (MLLMs) have demonstrated impressive performance in various vision-language (VL) tasks, but their expensive computations still limit the real-world…
Optimizing Few-Step Sampler for Diffusion Probabilistic Model
Jen-Yuan Huang
Diffusion Probabilistic Models (DPMs) have demonstrated exceptional capability of generating high-quality and diverse images, but their practical application is hindered by the int…