activity
20232026
most citedJoint Depth Prediction and Semantic Segmentation with Multi-View SAM

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens

Xinxuan Lu, Charless Fowlkes, Alexander C. Berg

Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global sc…

cs.CV2025

GViT: Representing Images as Gaussians for Visual Recognition

Jefferson Hernandez, Ruozhen He, Guha Balakrishnan +2

We introduce GVIT, a classification framework that abandons conventional pixel or patch grid input representations in favor of a compact set of learnable 2D Gaussians. Each image i…

cs.CV2024

Learning from Synthetic Data for Visual Grounding

Ruozhen He, Ziyan Yang, Paola Cascante-Bonilla +2

This paper extensively investigates the effectiveness of synthetic training data to improve the capabilities of vision-and-language models for grounding textual descriptions to ima…

cs.CV2024

Towards Artwork Explanation in Large-scale Vision Language Models

Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2

Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…

cs.CV2023

Improved Visual Grounding through Self-Consistent Explanations

Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang +2

Vision-and-language models trained to match images with text can be combined with visual explanation methods to point to the locations of specific objects in an image. Our work sho…

cs.CV20231 cited

Joint Depth Prediction and Semantic Segmentation with Multi-View SAM

Mykhailo Shvets, Dongxu Zhao, Marc Niethammer +2

Multi-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple v…