1 citations · 1 across the 3 of their papers we have counts for
6 papers · 1 filter
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
Xinxuan Lu, Charless Fowlkes, Alexander C. Berg
Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global sc…
GViT: Representing Images as Gaussians for Visual Recognition
Jefferson Hernandez, Ruozhen He, Guha Balakrishnan +2
We introduce GVIT, a classification framework that abandons conventional pixel or patch grid input representations in favor of a compact set of learnable 2D Gaussians. Each image i…
Learning from Synthetic Data for Visual Grounding
Ruozhen He, Ziyan Yang, Paola Cascante-Bonilla +2
This paper extensively investigates the effectiveness of synthetic training data to improve the capabilities of vision-and-language models for grounding textual descriptions to ima…
Towards Artwork Explanation in Large-scale Vision Language Models
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2
Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…
Improved Visual Grounding through Self-Consistent Explanations
Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang +2
Vision-and-language models trained to match images with text can be combined with visual explanation methods to point to the locations of specific objects in an image. Our work sho…
Joint Depth Prediction and Semantic Segmentation with Multi-View SAM
Mykhailo Shvets, Dongxu Zhao, Marc Niethammer +2
Multi-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple v…