#real-time inference
15 papers · 1 filter
Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles
Luca de Martino, Federico Aromolo, Federico Nesti +1
The paper proposes an efficient, real‑time anomaly segmentation pipeline for autonomous vehicles by reformulating PixOOD's Neyman‑Pearson scoring and deploying it with TensorRT, ac…
Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras
Edoardo Ragusa, Giovanni Paolo Canuti, Simone Lugani +2
The paper proposes hardware‑aware neural architecture search and a fine‑tuning pipeline to integrate depth data into compact RGB‑D networks for affordance segmentation on embedded…
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +7
TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…
WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing
Prathyush Sajith, Emadeldeen Hamdan, Ahmet Enis Cetin
The paper proposes replacing the global self‑attention stage in stereo transformer models with a data‑independent Walsh‑Hadamard token mixer, achieving similar depth accuracy while…
Stitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction
Shunsuke Kikuchi, Atsushi Kouno, Hiroki Matsuzaki
The paper introduces Stitch-Inferencer, a real‑time, model‑agnostic framework that builds an explicit panoramic canvas from successive endoscopic frames, allowing existing segmenta…
Reflex: Real-Time VLA Control through Streaming Inference
Yuanchun Guo, Bingyan Liu
Reflex is a framework that enables real-time streaming inference for vision‑language‑action models by exploiting a timestep‑invariance property to allow constant‑time attention cac…