◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Robik Shrestha

7 papers hereh-index 11908 citations15 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author4

Across the 7 of 7 papers where every author was matched, so the position is known.

fields
  • cs.CV3
  • cs.LG2
  • cs.AI1
  • cs.CL1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.CVShow all

3 papers · 1 filter

cs.CV2026

Visual Reasoning through Tool-supervised Reinforcement Learning

Qihua Dong, Gozde Sahin, Pei Wang +4

In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Models. To achieve that, we pr…

cs.CV2024

BloomVQA: Assessing Hierarchical Multi-modal Comprehension

Yunye Gong, Robik Shrestha, Jared Claypoole +4

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus…

cs.CV2024

Visual Grounding Methods for VQA are Working for the Wrong Reasons!

Robik Shrestha, Kushal Kafle, Christopher Kanan

Existing Visual Question Answering (VQA) methods tend to exploit dataset biases and spurious statistical correlations, instead of producing right answers for the right reasons. To…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.