11 papers
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +4
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as…
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson +5
Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different vis…
ShapeLib: Designing a library of programmatic 3D shape abstractions with Large Language Models
R. Kenny Jones, Paul Guerrero, Niloy J. Mitra +1
We present ShapeLib, the first method that uses the priors of Large Language Models (LLMs) to design libraries of programmatic 3D shape abstractions. Our system accepts two forms o…
ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes
Honglin Chen, Karran Pandey, Rundi Wu +6
Kinematic rigs provide a structured interface for articulating 3D meshes but lack any associated pose space, i.e., an explicit representation of the plausible manifold of joint con…
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra +3
Video generation has achieved remarkable progress in visual fidelity and controllability, enabling conditioning on text, layout, or motion. Among these, motion control - specifying…
LoST: Level of Semantics Tokenization for 3D Shapes
Niladri Shekhar Dutt, Zifan Shi, Paul Guerrero +4
Tokenization is a fundamental technique in the generative modeling of various modalities. In particular, it plays a critical role in autoregressive (AR) models, which have recently…