Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models
Yuchen Deng, Zidang Cai, Hai-Tao Zheng +3
Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video…
cs.AI2026
AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
Xinzi Cao, Jianyang Zhai, Pengfei Li +17
To meet the ever-increasing demand for computational efficiency, Neural Processing Units (NPUs) have become critical in modern AI infrastructure. However, unlocking their full pote…