18 papers
QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
Arian Kheirandish, Fardin Ayar, Ehsan Javanmardi +2
Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consistency through video-level s…
DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video
Muhammad Sulthan Adhipradhana, Ehsan Javanmardi, Naren Bao +1
Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a pedestrian accident is a fatal…
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
Yun Li, Jiachen Gong, Simon Thompson +7
Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing…
How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation
Saba Fathi, Fardin Ayar, Maryam Abdolali +3
Diffusion models have recently been repurposed for zero-shot classification, giving rise to diffusion classifiers that identify the best-matching text prompt by minimizing the nois…
Causal Scene Narration with Runtime Safety Supervision for Vision-Language-Action Driving
Yun Li, Yidu Zhang, Simon Thompson +2
Vision-Language-Action (VLA) models for autonomous driving must integrate diverse textual inputs, including navigation commands, hazard warnings, and traffic state descriptions, ye…
SUG-Occ: Explicit Semantics and Uncertainty Guided Sparse Learning for Efficient 3D Occupancy Prediction
Hanlin Wu, Pengfei Lin, Ehsan Javanmardi +4
3D semantic occupancy prediction has emerged as a critical perception task for autonomous driving due to its ability to offer voxel-level semantic and geometric understanding of th…