13 papers
ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation
Ruihang Zhang, Felix Taubner, Pooja Ravi +2
Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods…
REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception
Vincenzo Polizzi, David B. Lindell, Jonathan Kelly
Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and robustness to extreme lighting. However, exi…
VibES: Induced Vibration for Persistent Event-Based Sensing
Vincenzo Polizzi, Stephen Yang, Quentin Clark +3
Event cameras are a bio-inspired class of sensors that asynchronously measure per-pixel intensity changes. Under fixed illumination conditions in static or low-motion scenes, rigid…
Efficient and Training-Free Single-Image Diffusion Models
Haojun Qiu, Kiriakos N. Kutulakos, David B. Lindell
We consider the problem of generating images whose internal structure -- defined by the distribution of patches across multiple scales -- matches that of a single reference image.…
SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
Javad Rajabi, Kimia Shaban, Koorosh Roohi +2
Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resolutions beyond their training…
Velox: Learning Representations of 4D Geometry and Appearance
Anagh Malik, Dorian Chan, Xiaoming Zhao +3
We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downst…