◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jay Saraf

3 papers hereh-index 337 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI1
  • cs.CL1
  • cs.CV1

identity via Semantic Scholar / OpenAlex

most citedMM-PhyQA: Multimodal Physics Question-Answering With Multi-Image CoT Prompting

2 citations · 2 across the 2 of their papers we have counts for

collaborators

3 papers

cs.CV2025

Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning

Janak Kapuriya, Anwar Shaikh, Arnav Goel +8

In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLM…

cs.CL2024★ 2 cited

MM-PhyQA: Multimodal Physics Question-Answering With Multi-Image CoT Prompting

Avinash Anand, Janak Kapuriya, Apoorv Singh +5

While Large Language Models (LLMs) can achieve human-level performance in various tasks, they continue to face challenges when it comes to effectively tackling multi-step physics r…

cs.AI2024

MM-PhyRLHF: Reinforcement Learning Framework for Multimodal Physics Question-Answering

Janak Kapuriya, Chhavi Kirtani, Apoorv Singh +7

Recent advancements in LLMs have shown their significant potential in tasks like text summarization and generation. Yet, they often encounter difficulty while solving complex physi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.