14 citations · 14 across the 2 of their papers we have counts for
3 papers · 1 filter
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
Junyi Zhang, Charles Herrmann, Junhwa Hur +5
Estimating geometry from dynamic scenes, where objects move and deform over time, remains a core challenge in computer vision. Current approaches often rely on multi-stage pipeline…
Large Language Models are Visual Reasoning Coordinators
Liangyu Chen, Bo Li, Sheng Shen +5
Visual reasoning requires multimodal perception and commonsense cognition of the world. Recently, multiple vision-language models (VLMs) have been proposed with excellent commonsen…
CLAIR: Evaluating Image Captions with Large Language Models
David Chan, Suzanne Petryk, Joseph E. Gonzalez +2
The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, inc…