◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Hongbin Zhang

4 papers hereh-index 315 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.DC4
same name
  • Hongbin Zhang — 7 papers, h 3
  • Hongbin Zhang — 6 papers, h 2
  • Hongbin Zhang — 4 papers, h 2
  • Hongbin Zhang — 4 papers, h 3
  • Hongbin Zhang — 4 papers, h 2
  • Hongbin Zhang — 3 papers, h 6

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.DC2026

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System

Fengyao Bai, Hongbin Zhang, Zhitao Chen +3

High-throughput inference serving is essential for applications built on large language models (LLMs). Existing serving frameworks reduce request-level and batch-level bubbles thro…

cs.DC2026

PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers

Hongbin Zhang, Taosheng Wei, Jiazhi Jiang +3

Offline LLM inference seeks to maximize request processing under fixed budgets, making commodity GPU servers a promising choice. However, prior work typically considers offloading…

cs.DC2025

TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference

Hongbin Zhang, Taosheng Wei, Zhenyi Zheng +3

As the model size continuously increases, pipeline parallelism shows great promise in throughput-oriented LLM inference due to its low demand on communications. However, imbalanced…

cs.DC2025

EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration

Jiangsu Du, Hongbin Zhang, Taosheng Wei +4

Existing LLM serving strategies can be categorized based on whether prefill and decode phases are disaggregated: non-disaggregated (NoDG) or fully disaggregated (FuDG). However, th…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.