◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Somak Aditya

Indian Institute of Technology, Kharagpur

15 papers hereh-index 14743 citations40 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author4
  • middle author8
  • last author3

Across the 15 of 15 papers where every author was matched, so the position is known.

fields
  • cs.CL6
  • cs.AI4
  • cs.CV4
  • cs.LG1
affiliations
  • Indian Institute of Technology, Kharagpur
same name
  • Somak Aditya — 12 papers, h 4
  • Somak Aditya — 3 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20182025
most citedTrusting RoBERTa over BERT: Insights from CheckListing the Natural Language Inference Task

9 citations · 10 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2025

Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs

Sameep Vani, Shreyas Jena, Maitreya Patel +3

While Video Large Language Models (Video-LLMs) have demonstrated remarkable performance across general video understanding benchmarks-particularly in video captioning and descripti…

cs.CV2019

Integrating Knowledge and Reasoning in Image Understanding

Somak Aditya, Yezhou Yang, Chitta Baral

Deep learning based data-driven approaches have been successfully applied in various image understanding applications ranging from object recognition, semantic segmentation to visu…

cs.CV2018

Spatial Knowledge Distillation to aid Visual Reasoning

Somak Aditya, Rudra Saha, Yezhou Yang +1

For tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense)…

cs.CV2018

Explicit Reasoning over End-to-End Neural Architectures for Visual Question Answering

Somak Aditya, Yezhou Yang, Chitta Baral

Many vision and language tasks require commonsense reasoning beyond data-driven image and natural language processing. Here we adopt Visual Question Answering (VQA) as an example t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.