1 citations · 2 across the 14 of their papers we have counts for
6 papers · 1 filter
Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference
Ziming Dong, Hardik Sharma, Evan O'Toole +2
Large Language Models (LLMs) deliver state-of-the-art performance on complex reasoning tasks, but their inference costs limit deployment at scale. Small Language Models (SLMs) offe…
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
Vishnu Narayanan Moothedath, Umang Agarwal, Umeshraja N +3
We focus on a binary classification problem in an edge intelligence system where false negatives are more costly than false positives. The system has a compact, locally deployed mo…
Low-Regret and Low-Complexity Learning for Hierarchical Inference
Sameep Chattopadhyay, Vinay Sutar, Jaya Prakash Champati +1
This work focuses on Hierarchical Inference (HI) in edge intelligence systems, where a compact Local-ML model on an end-device works in conjunction with a high-accuracy Remote-ML m…
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
Adarsh Prasad Behera, Jaya Prakash Champati, Roberto Morabito +2
Recent progress in Language Models (LMs) has dramatically advanced the field of natural language processing (NLP), excelling at tasks like text generation, summarization, and quest…
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
Adarsh Prasad Behera, Paulius Daubaris, Iñaki Bravo +4
On-device inference holds great potential for increased energy efficiency, responsiveness, and privacy in edge ML systems. However, due to less capable ML models that can be embedd…
Online Algorithms for Hierarchical Inference in Deep Learning applications at the Edge
Vishnu Narayanan Moothedath, Jaya Prakash Champati, James Gross
We consider a resource-constrained Edge Device (ED), such as an IoT sensor or a microcontroller unit, embedded with a small-size ML model (S-ML) for a generic classification applic…