2 papers
cs.CL2026
ProbeLLM: Automating Principled Diagnosis of LLM Failures
Yue Huang, Zhengzhe Jiang, Yuchen Ma +8
Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has…
cs.CL2025
ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
Yue Huang, Zhengzhe Jiang, Xiaonan Luo +12
Empowering large language models (LLMs) with chemical intelligence remains a challenge due to the scarcity of high-quality, domain-specific instruction-response datasets and the mi…