collaborators

5 papers

cs.AR2025

D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations

Faraz Tahmasebi, Michael Pelluer, Hyoukjun Kwon

The computation and memory costs of large language models kept increasing over last decade, which reached over the scale of 1T parameters. To address the challenges from the large…

cs.LG2025

Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems

Rachid Karami, Rajeev Patwari, Hyoukjun Kwon +1

The integration of generative AI models, particularly large language models (LLMs), into real-time multi-model AI applications such as video conferencing and gaming is giving rise…

cs.AR2024

Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception

Mohanad Odema, Luke Chen, Hyoukjun Kwon +1

We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems…

cs.AR2024

FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI

Faraz Tahmasebi, Yian Wang, Benji Y. H. Huang +1

Recent research has shown that large language models (LLMs) can utilize low-precision floating point (FP) quantization to deliver high efficiency while maintaining original model a…

cs.CV2024

Efficient Depth Estimation for Unstable Stereo Camera Systems on AR Glasses

Yongfan Liu, Hyoukjun Kwon

Stereo depth estimation is a fundamental component in augmented reality (AR), which requires low latency for real-time processing. However, preprocessing such as rectification and…