11 papers
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Xinye Zhao, Spyridon Mastorakis
As large language models (LLMs) continue to scale, the memory footprint of key-value (KV) caches during inference has become a significant bottleneck. Existing approaches primarily…
An LLM-Powered AI Agent Framework for Holistic IoT Traffic Interpretation
Daniel Adu Worae, Spyridon Mastorakis
Internet of Things (IoT) networks generate diverse and high-volume traffic that reflects both normal activity and potential threats. Deriving meaningful insight from such telemetry…
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System
Sheikh Samit Muhaimin, Spyridon Mastorakis
The recent growth in the use of Large Language Models has made them vulnerable to sophisticated adversarial assaults, manipulative prompts, and encoded malicious inputs. Existing c…
I Know What You Did Last Summer: Identifying VR User Activity Through VR Network Traffic
Sheikh Samit Muhaimin, Spyridon Mastorakis
Virtual Reality (VR) technology has gained substantial traction and has the potential to transform a number of industries, including education, entertainment, and professional sect…
A Unified Framework for Context-Aware IoT Management and State-of-the-Art IoT Traffic Anomaly Detection
Daniel Adu Worae, Athar Sheikh, Spyridon Mastorakis
The rapid expansion of Internet of Things (IoT) ecosystems has introduced growing complexities in device management and network security. To address these challenges, we present a…
FusedInf: Efficient Swapping of DNN Models for On-Demand Serverless Inference Services on the Edge
Sifat Ut Taki, Arthi Padmanabhan, Spyridon Mastorakis
Edge AI computing boxes are a new class of computing devices that are aimed to revolutionize the AI industry. These compact and robust hardware units bring the power of AI processi…