2 citations · 5 across the 19 of their papers we have counts for
36 papers · 1 filter
A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources
Brunó B. Englert, Gijs Dubbelman
Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining…
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
Svetlana Orlova, Niccolò Cavagnero, Gijs Dubbelman
Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive video datasets, resulting in sub…
REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View
Kavin Chandrasekaran, Sorin Grigorescu, Gijs Dubbelman +1
A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar sensors, on the other hand, a…
Revisiting Radar Perception With Spectral Point Clouds
Hamza Alsharif, Jing Gu, Pavol Jancura +2
Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sparse point clouds, yet they…
Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models
Jing Gu, Niccolò Cavagnero, Gijs Dubbelman
Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems to handle rare and complex…
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Tommie Kerssies, Gabriele Berton, Ju He +5
Anticipating diverse future states is a central challenge in video world modeling. Discriminative world models produce a deterministic prediction that implicitly averages over poss…