1 citations · 1 across the 3 of their papers we have counts for
4 papers
ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement
Ryuji Oi, Hikari Otsuka, Kosuke Matsushima +4
Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow-matching-based VLA models have shown remarkabl…
Context Memorization for Efficient Long Context Generation
Yasuyuki Okoshi, Hao Mark Chen, Guanxi Lu +3
Modern large language model (LLM) applications increasingly rely on long conditioning prefixes to control model behavior at inference time. While prefix-augmented inference is effe…
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
Kosuke Matsushima, Yasuyuki Okoshi, Masato Motomura +1
Processing-in-Memory (PIM) architectures offer a promising solution to the memory bottlenecks in data-intensive machine learning, yet often overlook the growing challenge of activa…
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
Hao Mark Chen, Guanxi Lu, Yasuyuki Okoshi +3
Test-time scaling (TTS) has proven effective in enhancing the reasoning capabilities of large language models (LLMs). Verification plays a key role in TTS, simultaneously influenci…