collaborators

6 papers

cs.AI2026

Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

Mohanad Odema, Gabrielle De Micheli, Dayin Gou +3

Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-le…

cs.CL2025

3-Model Speculative Decoding

Sanghyun Byun, Mohanad Odema, Jung Ick Guack +3

Speculative Decoding (SD) accelerates inference in large language models by using a smaller draft model to propose tokens, which are then verified by a larger target model. However…

cs.CV2025

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema +3

Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and le…

cs.LG2025

CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression

Dayin Gou, Sanghyun Byun, Nilesh Malpeddi +4

Large Language Models (LLMs) typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. In particular, L…

cs.CL2025

APCE: Adaptive Progressive Context Expansion for Long Context Processing

Baisub Lee, Sanghyun Byun, Mohanad Odema +3

Deploying useful Long-Context Transformer Models (LCTMs) requires addressing two key challenges: (1) A growing memory footprint due to quadratic self-attention and linear KV-cache…

cs.CV2025

HunyuanVideo: A Systematic Framework For Large Video Generative Models

Weijie Kong, Qi Tian, Zijian Zhang +49

Recent advancements in video generation have significantly impacted daily life for both individuals and industries. However, the leading video generation models remain closed-sourc…