From the 2 of 24 linked papers with an AI index.
24 papers
ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM
Hyunwoo Oh, Suyeon Jang, Hanning Chen +3
Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs. Practical deployment spans fragmented regimes-f…
PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
Hyunwoo Oh, Suyeon Jang, Hanning Chen +4
PolyQ is a co-designed compiler and quantization framework that assigns per‑channel bit‑widths to LLM activations on CPUs, enabling fine‑grained fractional‑bit precision while keep…
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
Hyunwoo Oh, Hanning Chen, Sanggeon Yun +5
Multimodal stacks that mix ViTs, CNNs, GNNs, and transformer NLP strain embedded platforms because their compute/memory patterns diverge and hard real-time targets leave little sla…
TorR: Towards Brain-Inspired Task-Oriented Reasoning via Cache-Oriented Algorithm-Architecture Co-design
Hyunwoo Oh, SungHeon Jeong, Suyeon Jang +4
Task-oriented object detection (TOOD) atop CLIP offers open-vocabulary, prompt-driven semantics, yet dense per-window computation and heavy memory traffic hinder real-time, power-l…
MERIT: Multi-domain Efficient RAW Image Translation
Wenjun Huang, Shenghao Fu, Yian Jin +10
RAW images captured by different camera sensors exhibit substantial domain shifts due to varying spectral responses, noise characteristics, and tone behaviors, complicating their d…
Draft and Refine with Visual Experts
Sungheon Jeong, Ryozo Masukawa, Jihong Park +5
While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…