◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Panda

2 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author2

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.LG1
  • cs.PF1

identity via Semantic Scholar / OpenAlex

most citedPerformance Characterization of using Quantization for DNN Inference on Edge Devices: Extended Version

2 citations · 3 across the 2 of their papers we have counts for

collaborators

2 papers

cs.LG2024★ 1 cited

Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference

Jinghan Yao, Quentin Anthony, Aamir Shafi +3

In large language models like the Generative Pre-trained Transformer, the Mixture of Experts paradigm has emerged as a powerful technique for enhancing model expressiveness and acc…

cs.PF2023★ 2 cited

Performance Characterization of using Quantization for DNN Inference on Edge Devices: Extended Version

Hyunho Ahn, Tian Chen, Nawras Alnaasan +5

Quantization is a popular technique used in Deep Neural Networks (DNN) inference to reduce the size of models and improve the overall numerical performance by exploiting native har…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.