◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Wenjie Huang

5 papers hereh-index 212 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author5

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.DC3
  • cs.AI2
same name
  • Wenjie Huang — 11 papers, h 9
  • Wenjie Huang — 6 papers, h 4
  • Wenjie Huang — 5 papers, h 2
  • Wenjie Huang — 2 papers, h 3
  • Wenjie Huang — 2 papers, h 1
  • Wenjie Huang — 2 papers, h 1

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.AI2026

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Tianqi Xu, Lu Lv, Haoyang Huang +15

Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evalua…

cs.AI2026

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Haoyang Huang, Wenjie Huang, Tianqi Xu +14

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory…

cs.DC2025

High-Throughput LLM inference on Heterogeneous Clusters

Yi Xiong, Jinqi Huang, Wenjie Huang +6

Nowadays, many companies possess various types of AI accelerators, forming heterogeneous clusters. Efficiently leveraging these clusters for high-throughput large language model (L…

cs.DC2025

SLO-Aware Scheduling for Large Language Model Inferences

Jinqi Huang, Yi Xiong, Xuebing Yu +4

Large language models (LLMs) have revolutionized applications such as code completion, chatbots, and online classification. To elevate user experiences, service level objectives (S…

cs.DC2025

WindVE: Collaborative CPU-NPU Vector Embedding

Jinqi Huang, Xuebing Yu, Yi Xiong +4

Retrieval-Augmented Generation is a technology that enhances large language models by integrating information retrieval. In the industry, inference services based on LLMs are highl…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.