3 papers
cs.CV2025
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
Weijie Shen, Xinrui Wang, Yuanqi Nie +1
Current Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) excel in single-turn tasks but face significant challenges in multi-turn interactions requiring deep c…
cs.CV2025
Revolutionizing Radiology Workflow with Factual and Efficient CXR Report Generation
Pimchanok Sukjai, Apiradee Boonmee
The escalating demand for medical image interpretation underscores the critical need for advanced artificial intelligence solutions to enhance the efficiency and accuracy of radiol…
cs.CV2025
Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks
Leo Franklin, Apiradee Boonmee, Kritsada Wongsuwan
Vision generation remains a challenging frontier in artificial intelligence, requiring seamless integration of visual understanding and generative capabilities. In this paper, we p…