activity
20182026
most citedTowards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2026

Speculation at a Distance: Where Edge-Cloud Speculative Decoding Actually Pays Off

Yuan Lyu, Bharath Irukulapati, Jaya Prakash Champati

Speculative decoding (SD) accelerates LLM inference by - times when the draft and target models are co-located. This has motivated a distributed variant (DSD) that places t…

cs.DC2024

Error Bounds for the Network Scale-Up Method

Sergio Díaz-Aranda, Juan Marcos Ramírez, Mohit Daga +4

Epidemiologists and social scientists have used the Network Scale-Up Method (NSUM) for over thirty years to estimate the size of a hidden sub-population within a social network. Th…

cs.DC2024

Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices

Adarsh Prasad Behera, Roberto Morabito, Joerg Widmer +1

The Hierarchical Inference (HI) paradigm employs a tiered processing: the inference from simple data samples are accepted at the end device, while complex data samples are offloade…

cs.DC2023

The Case for Hierarchical Deep Learning Inference at the Network Edge

Ghina Al-Atat, Andrea Fresa, Adarsh Prasad Behera +3

Resource-constrained Edge Devices (EDs), e.g., IoT sensors and microcontroller units, are expected to make intelligent decisions using Deep Learning (DL) inference at the edge of t…

cs.DC2022

Edge-MultiAI: Multi-Tenancy of Latency-Sensitive Deep Learning Applications on Edge

SM Zobaed, Ali Mokhtari, Jaya Prakash Champati +2

Smart IoT-based systems often desire continuous execution of multiple latency-sensitive Deep Learning (DL) applications. The edge servers serve as the cornerstone of such IoT-based…