Publications (8)
Palu: Compressing KV-Cache with Low-Rank Projection
Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin +7
Post-training KV-Cache compression methods typically either sample a subset of effectual tokens or quantize the data into lower numerical bit width. However, these methods cannot e…
The MSP-Podcast Corpus
Carlos Busso, Reza Lotfian, Kusha Sridhar +9
The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing datab…
GazeNLQ @ Ego4D Natural Language Queries Challenge 2025
Wei-Cheng Lin, Chih-Ming Lien, Chen Lo +1
This report presents our solution to the Ego4D Natural Language Queries (NLQ) Challenge at CVPR 2025. Egocentric video captures the scene from the wearer's perspective, where gaze…
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
Francesca Ronchini, Ho-Hsiang Wu, Wei-Cheng Lin +1
This paper investigates the design of effective prompt strategies for generating realistic datasets using Text-To-Audio (TTA) models. We also analyze different techniques for effic…
Q-YOLOP: Quantization-aware You Only Look Once for Panoptic Driving Perception
Chi-Chih Chang, Wei-Cheng Lin, Pei-Shuo Wang +4
In this work, we present an efficient and quantization-aware panoptic driving perception model (Q- YOLOP) for object detection, drivable area segmentation, and lane line segmentati…
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin +8
Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent studies attempt to share KV-Cache…
ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration
Ning-Chi Huang, Chi-Chih Chang, Wei-Cheng Lin +3
sparsity is an emerging model compression method supported by more and more accelerators to speed up sparse matrix multiplication in deep neural networks. Most existing $N{…
Versatile audio-visual learning for emotion recognition
Lucas Goncalves, Seong-Gyun Leem, Wei-Cheng Lin +2
Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only…