2 papers
cs.CV2025
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
Janak Kapuriya, Anwar Shaikh, Arnav Goel +8
In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLM…
cs.AI2025
MM-PhyRLHF: Reinforcement Learning Framework for Multimodal Physics Question-Answering
Janak Kapuriya, Chhavi Kirtani, Apoorv Singh +7
Recent advancements in LLMs have shown their significant potential in tasks like text summarization and generation. Yet, they often encounter difficulty while solving complex physi…