activity
20242026
most citedGemma 4 Technical Report

1 citations · 2 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL20261 cited

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…

cs.CV2026

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

Gabriel Fiastre, Antoine Yang, Cordelia Schmid

Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal…

cs.CV2026

Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

Shoubin Yu, Lei Shu, Antoine Yang +6

Multimodal AI agents are increasingly automating complex real-world workflows that involve online web execution. However, current web-agent benchmarks suffer from a critical limita…

cs.CV2025

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

Lucas Ventura, Antoine Yang, Cordelia Schmid +1

We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, a…

cs.CL2025

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +209

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision underst…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…