From the 1 of 14 linked papers with an AI index.
14 papers
TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs
Jin Zhang, Linyu Li, Weili Jiang +7
Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions…
Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
Linyu Li, Zhi Jin, Yichi Zhang +6
The paper introduces EC-Reason-Bench, a training-free diagnostic benchmark that evaluates why general large language models struggle with detailed enzyme classification and how per…
From Chat to Interview: Agentic Requirements Elicitation with an Experience Ontology
Dongming Jin, Zhi Jin, Yaotian Yang +5
Requirements elicitation interviews are crucial and time-consuming in requirements engineering, but heavily rely on the experience of requirements analysts. Although recent advance…
SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
Xingyan Liu, Xiyue Luo, Linyu Li +3
Deploying LLM-powered agents in enterprise scenarios such as cloud technical support demands high-quality, domain-specific skills. However, existing skill creators lack domain grou…
When Modalities Remember: Continual Learning for Multimodal Knowledge Graphs
Linyu Li, Zhi Jin, Yichi Zhang +5
Real-world multimodal knowledge graphs (MMKGs) are dynamic, with new entities, relations, and multimodal knowledge emerging over time. Existing continual knowledge graph reasoning…
ReqElicitGym: An Evaluation Environment for Interview Competence in Conversational Requirements Elicitation
Dongming Jin, Zhi Jin, Zheng Fang +4
With the rapid improvement of LLMs' coding capabilities, the bottleneck of LLM-based automated software development is shifting from generating correct code to eliciting users' req…