5 papers
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
Jinhao Jing, Zheng Ma, Jinwei Liang +9
Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address thi…
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
Ben Yao, Qiuchi Li, Yazhou Zhang +4
While LLMs have demonstrated medical knowledge and conversational ability, their deployment in clinical practice raises new risks: patients may place greater trust in LLM-generated…
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
Wenya Xie, Qingying Xiao, Yu Zheng +8
The rise of large language models (LLMs) has transformed healthcare by offering clinical guidance, yet their direct deployment to patients poses safety risks due to limited domain…
Is Your LLM Outdated? A Deep Look at Temporal Generalization
Chenghao Zhu, Nuo Chen, Yufei Gao +3
The rapid advancement of Large Language Models (LLMs) has led to the development of benchmarks that consider temporal dynamics, however, there remains a gap in understanding how we…
Mixture of Latent Experts Using Tensor Products
Zhan Su, Fengran Mo, Prayag Tiwari +3
In multi-task learning, the conventional approach involves training a model on multiple tasks simultaneously. However, the training signals from different tasks can interfere with…