2 papers
cs.LG2026
DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation
Tan T. Nguyen, Quan V. Dang
As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually b…
cs.CL2026
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
Dang Quang Thien Tran, Quang V. Dang, Vinamra Tyagi +7
As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety. While unimodal…