4 papers
ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation
Lili Gao, Yanbo Xu, William Koch +8
We introduce ScenarioControl, the first vision-language control mechanism for learned driving scenario generation. Given a text prompt or an input image, Scenario-Control synthesiz…
Telescope: Learnable Hyperbolic Foveation for Ultra-Long-Range Object Detection
Parker Ewen, Dmitriy Rivkin, Mario Bijelic +1
Autonomous highway driving, especially for long-haul heavy trucks, requires detecting objects at long ranges beyond 500 meters to satisfy braking distance requirements at high spee…
ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation
Dmitriy Rivkin, Parker Ewen, Lili Gao +5
Recent video diffusion models achieve high-quality generation through recurrent frame processing where each frame generation depends on previous frames. However, this recurrent mec…
PhotoBot: Reference-Guided Interactive Photography via Natural Language
Oliver Limoyo, Jimmy Li, Dmitriy Rivkin +2
We introduce PhotoBot, a framework for fully automated photo acquisition based on an interplay between high-level human language guidance and a robot photographer. We propose to co…