2 papers
cs.CL2026
Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion
Zhen Cheng, Hao-Bo Yang, Wan-Yi Huang +1
Key-Value (KV) cache memory and bandwidth increasingly dominate large language model inference cost in long-context and long-generation regimes. Architectures such as multi-head la…
cs.AI2025
Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
Zhenliang Zhang, Yuxi Wang, Hongzhao Xie +6
A key feature differentiating artificial general intelligence (AGI) from traditional AI is that AGI can perform composite tasks that require a wide range of capabilities. Although…