7 papers · 1 filter
Dissecting Failure Dynamics in Large Language Model Reasoning
Wei Zhu, Jian Zhang, Lixing Yu +2
Large Language Models (LLMs) achieve strong performance through extended inference-time deliberation, yet how their reasoning failures arise remains poorly understood. By analyzing…
Toward Clinically Explainable AI for Medical Diagnosis: A Foundation Model with Human-Compatible Reasoning via Reinforcement Learning
Qika Lin, Yifan Zhu, Bin Pu +14
The clinical adoption of artificial intelligence (AI) in medical diagnostics is critically hampered by its black-box nature, which prevents clinicians from verifying the rationale…
Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction Refinement
Jian Zhang, Zhangqi Wang, Zhiyuan Wang +5
Linguistic expressions of emotions such as depression, anxiety, and trauma-related states are pervasive in clinical notes, counseling dialogues, and online mental health communitie…
ErrEval: Error-Aware Evaluation for Question Generation through Explicit Diagnostics
Weiping Fu, Bifan Wei, Jingyi Hao +7
Automatic Question Generation (QG) often produces outputs with critical defects, such as factual hallucinations and answer mismatches. However, existing evaluation methods, includi…
-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation
Jian Zhang, Yu He, Zhiyuan Wang +5
Scientific reasoning relies not only on logical inference but also on activating prior knowledge and experiential structures. Memory can efficiently reuse knowledge and enhance rea…
MAXS: Meta-Adaptive Exploration with LLM Agents
Jian Zhang, Zhiyuan Wang, Zhangqi Wang +7
Large Language Model (LLM) Agents exhibit inherent reasoning abilities through the collaboration of multiple tools. However, during agent inference, existing methods often suffer f…