◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Thomas Kollar

12 papers hereh-index 103.6k citations17 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author7
  • last author2

Across the 9 of 12 papers where every author was matched, so the position is known.

fields
  • cs.RO6
  • cs.CV3
  • cs.CL2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CVShow all

3 papers · 1 filter

cs.CV2025

Understanding Complexity in VideoQA via Visual Program Generation

Cristobal Eyzaguirre, Igor Vasiljevic, Achal Dave +5

We propose a data-driven approach to analyzing query complexity in Video Question Answering (VideoQA). Previous efforts in benchmark design have relied on human expertise to design…

cs.CV2025

Should VLMs be Pre-trained with Image Data?

Sedrick Keh, Jean Mercat, Samir Yitzhak Gadre +8

Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase effectively unlocks this capabil…

cs.CV2024

Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna +3

Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fu…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.