4 papers
CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
Xue-Jian Gao, Deng Pan, Yueming Su +12
AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost e…
Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference
Lingchao Zheng, Yuwei Fan, Jun Li +5
Quantization is essential for efficient large language model (LLM) inference, yet the dequantization step-converting low-bit weights back to high-precision for matrix multiplicatio…
Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
Haotong Qin, Cheng Hu, Michele Magno
Large Language Model (LLM)-based Vision-Language Models (VLMs) have substantially extended the boundaries of visual understanding capabilities. However, their high computational de…
Drive Fast, Learn Faster: On-Board RL for High Performance Autonomous Racing
Benedict Hildisch, Edoardo Ghignone, Nicolas Baumann +3
Autonomous racing presents unique challenges due to its non-linear dynamics, the high speed involved, and the critical need for real-time decision-making under dynamic and unpredic…