3 papers
cs.CL2026
ProbeLLM: Automating Principled Diagnosis of LLM Failures
Yue Huang, Zhengzhe Jiang, Yuchen Ma +8
Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has…
cs.SE2026
CODEFUSE-DEBENCH: An Empirical Study on Readability, Recompilability, and Functionality
Puzhuo Liu, Yuhan Huang, Jianlei Chi +2
Binary decompilation aims to recover binaries into high-level source code, but existing evaluations mainly rely on syntactic similarity or single-axis readability metrics, which fa…
cs.AR2024
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
Jingchen Zhu, Chenhao Xue, Yiqi Chen +13
The emergence of the large language model~(LLM) poses an exponential growth of demand for computation throughput, memory capacity, and communication bandwidth. Such a demand growth…