◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Haifeng Wu

4 papers hereh-index 27 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1
  • last author1

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1
same name
  • Haifeng Wu — 2 papers, h 2
  • Haifeng Wu — 2 papers, h 1
  • Haifeng Wu — 1 paper, h 0

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2026

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

Srinivasan Manoharan, Junhua Zhao, Fangbo Tu +6

Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries, escalations, and developer…

cs.LG2026

RLM-Cascade: Response-Level Speculative Decoding for Cost-Efficient LLM API Serving

Haifeng Wu, Srinivasan Manoharan, Fangbo Tu +2

We present RLM-Cascade, a proxy-layer system that applies speculative decoding at the response level to reduce LLM API costs without requiring model architecture access or a shared…

cs.LG2026

Domain-Adapted Small Language Models with Hybrid Post-Processing: Achieving Cost-Efficient, Low-Latency Multi-Label Structured Prediction via LoRA Fine-Tuning on Scarce Data

Srinivasan Manoharan, Dilipkumar Nallusamy, Sachin Kumar +1

Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks incurs prohibitive latency, cost, and data-privacy overhead. We present a hybrid fra…

cs.AI2026

Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation

Fangbo Tu, Junhua Zhao, Chi Liu +4

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production enviro…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.