◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Sergi Caelles

3 papers hereh-index 1310.9k citations27 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1

Across the 1 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CV2
  • cs.CL1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.CV2024

VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models

Harshit, Tolga Tasdizen

The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Mode…

cs.CV2024

Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers

Aleksandar Stanić, Sergi Caelles, Michael Tschannen

Visual reasoning is dominated by end-to-end neural networks scaled to billions of model parameters and training examples. However, even the largest models struggle with composition…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.