Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ProtStructQA: A Denotation Threshold in Protein Structural Reasoning
Aravind Mandiga, Guoming Li, Jin Lu +3
Protein-language systems are often evaluated by whether they generate plausible biological text, but a structural question has a sharper semantics: it denotes a measurement in a 3D…
cs.CL2025
Fact or Guesswork? Evaluating Large Language Models' Medical Knowledge with Structured One-Hop Judgments
Jiaxi Li, Yiwei Wang, Kai Zhang +5
Large language models (LLMs) have been widely adopted in various downstream task domains. However, their abilities to directly recall and apply factual medical knowledge remains un…