2 papers
cs.CV2026
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
Xinxuan Lu, Charless Fowlkes, Alexander C. Berg
Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global sc…
cs.CV2025
GViT: Representing Images as Gaussians for Visual Recognition
Jefferson Hernandez, Ruozhen He, Guha Balakrishnan +2
We introduce GVIT, a classification framework that abandons conventional pixel or patch grid input representations in favor of a compact set of learnable 2D Gaussians. Each image i…