3 papers
cs.CL2026
Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models
Amogh Sheth, Biruk Assefa, Yi Wen Huang +2
Large language models (LLMs) excel at multi-step reasoning but incur substantial inference cost. We introduce Causal Attribution Pruning (CAP), a training-free method that identifi…
cs.DC2026
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
Zhongzhen Wen, Shudi Shao, Zhong Li +4
The performance of deep learning models critically depends on efficient kernel implementations, yet developing high-performance kernels for specialized accelerators remains time-co…
cs.SE2025
Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
Yu Ge, Linna Xie, Zhong Li +2
Large Language Model Powered Multi-Agent Systems (MASs) are increasingly employed to automate complex real-world problems, such as programming and scientific discovery. Despite the…