2 papers
cs.LG2026
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
Ahmed Šabanović, Paul Joe Maliakel, Ivona Brandić
Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but incurs communication delay and…
cs.LG2025
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
Paul Joe Maliakel, Shashikant Ilager, Ivona Brandic
LLM inference exhibits substantial variability across queries and execution phases, yet inference configurations are often applied uniformly. We present a measurement-driven charac…