4 papers
StructKV: Preserving the Structural Skeleton for Scalable Long-Context Inference
Zhirui Chen, Peiyang Liu, Ling Shao
As Large Language Models (LLMs) scale to support context windows exceeding one million tokens, the linear growth of Key-Value (KV) cache imposes severe memory capacity and bandwidt…
Self-Consolidation for Self-Evolving Agents
Hongzhuo Yu, Fei Zhu, Guo-Sen Xie +1
While large language model (LLM) agents have demonstrated impressive problem-solving capabilities, they typically operate as static systems, lacking the ability to evolve through l…
MetaTPT: Meta Test-time Prompt Tuning for Vision-Language Models
Yuqing Lei, Yingjun Du, Yawen Huang +2
Vision-language models (VLMs) such as CLIP exhibit strong zero-shot generalization but remain sensitive to domain shifts at test time. Test-time prompt tuning (TPT) mitigates this…
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
Wenbo Lyu, Yingjun Du, Jinglin Zhao +2
Understanding multi-image, multi-turn scenarios is a critical yet underexplored capability for Large Vision-Language Models (LVLMs). Existing benchmarks predominantly focus on stat…