From the 2 of 31 linked papers with an AI index.
7 papers · 1 filter
ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM
Hyunwoo Oh, Suyeon Jang, Hanning Chen +3
Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs. Practical deployment spans fragmented regimes-f…
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
Hyunwoo Oh, Hanning Chen, Sanggeon Yun +5
Multimodal stacks that mix ViTs, CNNs, GNNs, and transformer NLP strain embedded platforms because their compute/memory patterns diverge and hard real-time targets leave little sla…
TorR: Towards Brain-Inspired Task-Oriented Reasoning via Cache-Oriented Algorithm-Architecture Co-design
Hyunwoo Oh, SungHeon Jeong, Suyeon Jang +4
Task-oriented object detection (TOOD) atop CLIP offers open-vocabulary, prompt-driven semantics, yet dense per-window computation and heavy memory traffic hinder real-time, power-l…
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
Hyunwoo Oh, Hanning Chen, Sanggeon Yun +5
Deformable transformers deliver state-of-the-art detection but map poorly to hardware due to irregular memory access and low arithmetic intensity. We introduce QUILL, a schedule-aw…
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
Hyunwoo Oh, KyungIn Nam, Rajat Bhattacharjya +7
Recent advances in LLMs have outpaced the computational and memory capacities of edge platforms that primarily employ CPUs, thereby challenging efficient and scalable deployment. W…
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
Tamoghno Das, Khanh Phan Vu, Hanning Chen +2
The integration of spiking neural networks (SNNs) with transformer-based architectures has opened new opportunities for bio-inspired low-power, event-driven visual reasoning on edg…