4 papers
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
You-Liang Huang, Xinhao Huang, Chengxi Liao +1
Existing works on large language model (LLM) decomposition mainly focus on improving performance on downstream tasks, but they ignore the poor parallel inference performance when t…
SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model Compression
Xinhao Huang, You-Liang Huang, Zeyi Wen
Large language models (LLMs) have demonstrated impressive capabilities across various tasks, but the billion-scale parameters pose deployment challenges. Although existing methods…
DOne: Decoupling Structure and Rendering for High-Fidelity Design-to-Code Generation
Xinhao Huang, Jinke Yu, Wenhao Xu +5
While Vision Language Models (VLMs) have shown promise in Design-to-Code generation, they suffer from a "holistic bottleneck-failing to reconcile high-level structural hierarchy wi…
SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
Xinhao Huang, Zhibo Ren, Yipeng Yu +3
In long structured document retrieval, existing methods typically fine-tune pre-trained language models (PLMs) using contrastive learning on datasets lacking explicit structural in…