2 papers
cs.CV2026
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
Qirui Jiao, Daoyuan Chen, Yilun Huang +3
While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the long, detailed prompts required for prof…
cs.CL2024
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
Jiayi Kuang, Jingyou Xie, Haohao Luo +6
Visual Question Answering (VQA) is a challenge task that combines natural language processing and computer vision techniques and gradually becomes a benchmark test task in multimod…