9 papers
Human Insights Driven Latent Space for Different Driving Perspectives: A Unified Encoder for Efficient Multi-Task Inference
Huy-Dung Nguyen, Anass Bairouk, Mirjana Maras +6
Autonomous driving systems require a comprehensive understanding of the environment, achieved by extracting visual features essential for perception, planning, and control. However…
See Less, Drive Better: Generalizable End-to-End Autonomous Driving via Foundation Models Stochastic Patch Selection
Amir Mallak, Erfan Aasi, Shiva Sreeram +3
Recent advances in end-to-end autonomous driving show that policies trained on patch-aligned features extracted from foundation models generalize better to Out-of-Distribution (OOD…
ReGen: Generative Robot Simulation via Inverse Design
Phat Nguyen, Tsun-Hsuan Wang, Zhang-Wei Hong +5
Simulation plays a key role in scaling robot learning and validating policies, but constructing simulations remains a labor-intensive process. This paper introduces ReGen, a genera…
Holistic Surgical Phase Recognition with Hierarchical Input Dependent State Space Models
Haoyang Wu, Tsun-Hsuan Wang, Mathias Lechner +7
Surgical workflow analysis is essential in robot-assisted surgeries, yet the long duration of such procedures poses significant challenges for comprehensive video analysis. Recent…
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
Makram Chahine, Alex Quach, Alaa Maalouf +2
End-to-end learning directly maps sensory inputs to actions, creating highly integrated and efficient policies for complex robotics tasks. However, such models often struggle to ge…
DataS^3: Dataset Subset Selection for Specialization
Neha Hulkund, Alaa Maalouf, Levi Cai +15
In many real-world machine learning (ML) applications (e.g. detecting broken bones in x-ray images, detecting species in camera traps), in practice models need to perform well on s…