3 papers
cs.LG2026
Value-Aware Stochastic KV Cache Eviction for Reasoning Models
Ting-Yun Chang, Harvey Yiyun Fu, Deqing Fu +3
Reasoning models improve accuracy through extended chains of thought, but their long outputs create a memory and compute bottleneck. KV cache eviction methods reduce this cost by e…
cs.LG2025
Why Do Some Inputs Break Low-Bit LLM Quantization?
Ting-Yun Chang, Muru Zhang, Jesse Thomason +1
Low-bit weight-only quantization significantly reduces the memory footprint of large language models (LLMs), but disproportionately affects certain examples. We analyze diverse 3-4…
cs.RO2025
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
Wang Bill Zhu, Miaosen Chai, Ishika Singh +2
We propose PSALM-V, the first autonomous neuro-symbolic learning system able to induce symbolic action semantics (i.e., pre- and post-conditions) in visual environments through int…