3 papers
cs.AR2026
Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding
Soongyu Choi, Yuntae Kim, Muyoung Son +1
Speculative decoding has emerged as a promising lossless approach for accelerating Large Language Models (LLMs). As reasoning LLMs increasingly suffer from decode-stage overhead an…
cs.LG2026
CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution
Muyoung Son, Yi Chen, Seungjae Yoo +2
The Mixture-of-Experts (MoE) architecture improves computational efficiency via sparse expert activation, but throughput-oriented inference faces substantial GPU memory pressure du…
cs.CV2025
AoP-SAM: Automation of Prompts for Efficient Segmentation
Yi Chen, Mu-Young Son, Chuanbo Hua +1
The Segment Anything Model (SAM) is a powerful foundation model for image segmentation, showing robust zero-shot generalization through prompt engineering. However, relying on manu…