29 papers
Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding
Duo Zhang, Zhehui Yin, Zhiyun Yao +8
Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and contactless sensing modality, b…
FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs
Jiawei Liang, Haotong Qin, Linfeng Du +7
Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications. While Field-Programmable…
PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation
Pietro Bonazzi, Nicola Farronato, Stefan Zihlmann +2
Real-time, on-device segmentation is critical for latency-sensitive and privacy-aware applications such as smart glasses and Internet-of-Things devices. We introduce PicoSAM3, a li…
SED:Lightweight Saliency prediction for Event-based data via Distillation
Romaric Mazna, Jean Martinet, Michele Magno
Event-based saliency prediction has gained attention recently, as combining event cameras with saliency estimation can act as an upstream stage that naturally improves the efficien…
WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
Weilun Feng, Guoxin Fan, Haotong Qin +10
Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts.…
Fast-SAM3D: 3Dfy Anything in Images but Faster
Weilun Feng, Mingqiang Wu, Zhiliang Chen +10
SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the \textbf{firs…