4 papers
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
Xinxuan Lu, Charless Fowlkes, Alexander C. Berg
Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global sc…
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions
Kyungmin Kim, JB Lanier, Pierre Baldi +2
Recent advancements in Model-Based Reinforcement Learning (MBRL) have made it a powerful tool for visual control tasks. Despite improved data efficiency, it remains challenging to…
CriSp: Leveraging Tread Depth Maps for Enhanced Crime-Scene Shoeprint Matching
Samia Shafique, Shu Kong, Charless Fowlkes
Shoeprints are a common type of evidence found at crime scenes and are used regularly in forensic investigations. However, existing methods cannot effectively employ deep learning…
Instance Tracking in 3D Scenes from Egocentric Videos
Yunhan Zhao, Haoyu Ma, Shu Kong +1
Egocentric sensors such as AR/VR devices capture human-object interactions and offer the potential to provide task-assistance by recalling 3D locations of objects of interest in th…