3 papers
cs.AR2026
Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels
Amir Taherin, Sana Taghipour Anvari, Charles Amante +9
Edge LLM deployment is shaped by more than model size and precision: inference backend, hardware platform, memory traffic, and power management all affect latency and efficiency. W…
cs.AI2026
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG
Zlatan Feric, Amir Taherin, Yanzhi Wang +1
Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prom…
cs.AI2025
Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
Amir Taherin, Juyi Lin, Arash Akbari +5
Vision-Language-Action (VLA) models have emerged as powerful generalist policies for robotic control, yet their performance scaling across model architectures and hardware platform…