most citedAQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization

1 citations · 1 across the 3 of their papers we have counts for

collaborators

8 papers

cs.RO2026

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

Ryuji Oi, Hikari Otsuka, Kosuke Matsushima +4

Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow-matching-based VLA models have shown remarkabl…

cs.CL2026

Context Memorization for Efficient Long Context Generation

Yasuyuki Okoshi, Hao Mark Chen, Guanxi Lu +3

Modern large language model (LLM) applications increasingly rely on long conditioning prefixes to control model behavior at inference time. While prefix-augmented inference is effe…

cs.AR20261 cited

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization

Kosuke Matsushima, Yasuyuki Okoshi, Masato Motomura +1

Processing-in-Memory (PIM) architectures offer a promising solution to the memory bottlenecks in data-intensive machine learning, yet often overlook the growing challenge of activa…

cs.LG2026

AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size

Guanxi Lu, Hao Mark Chen, Yuto Karashima +3

Diffusion-based large language models (dLLMs) are gaining attention for their inherent capacity for parallel decoding, offering a compelling alternative to autoregressive LLMs. Amo…

cs.LG2025

The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms

Hikari Otsuka, Daiki Chijiwa, Yasuyuki Okoshi +3

The strong lottery ticket hypothesis (SLTH) conjectures that high-performing subnetworks, called strong lottery tickets (SLTs), are hidden in randomly initialized neural networks.…

cs.AR2025

DX100: A Programmable Data Access Accelerator for Indirection

Alireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu +8

Indirect memory accesses frequently appear in applications where memory bandwidth is a critical bottleneck. Prior indirect memory access proposals, such as indirect prefetchers, ru…