3 papers
cs.CL2025
Coherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language Models
Wenjie Luo, Ruocheng Li, Shanshan Zhu +1
Despite significant advancements, current large language models (LLMs) and vision-language models (LVLMs) continue to struggle with complex, multi-step, cross-modal common sense re…
cs.CL2025
Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models
Julian Perry, Frank Sanders, Carter Scott
In this paper, we presents a novel method for improving text-to-image generation by combining Large Language Models (LLMs) with diffusion models, a hybrid approach aimed at achievi…
cs.CL2025
Dynamic Knowledge Integration for Enhanced Vision-Language Reasoning
Julian Perry, Surasakdi Siripong, Thanakorn Phonchai
Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal tasks, but their performance is often constrained by the lack of external knowledge int…