2 papers
cs.CV2026
Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere
Jiayang He, Tianling Xu, Diancheng Kang +4
Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent of the input. These tokens of…
cs.CL2026
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
Diancheng Kang, Zheyuan Liu, Ningshan Ma +3
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful…