activity
20242026
collaborators

5 papers

cs.CL2026

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding

WenHung Lee, Jian-Jia Chen, Xiaolin Lin +6

While speculative decoding improves inference throughput for multi-batch long-context Large Language Models (LLMs), its efficiency is often limited by a verification bottleneck whe…

cs.CL2025

Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding

Pei-Shuo Wang, Jian-Jia Chen, Chun-Che Yang +4

The immense model sizes of large language models (LLMs) challenge deployment on memory-limited consumer GPUs. Although model compression and parameter offloading are common strateg…

cs.AR2025

Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration

Endri Taka, Ning-Chi Huang, Chi-Chih Chang +3

FPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic rese…

cs.CV2024

V"Mean"ba: Visual State Space Models only need 1 hidden dimension

Tien-Yu Chi, Hung-Yueh Chiang, Chi-Chih Chang +2

Vision transformers dominate image processing tasks due to their superior performance. However, the quadratic complexity of self-attention limits the scalability of these systems a…

cs.AI2024

Palu: Compressing KV-Cache with Low-Rank Projection

Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin +7

Post-training KV-Cache compression methods typically either sample a subset of effectual tokens or quantize the data into lower numerical bit width. However, these methods cannot e…