4 papers
From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions
Xiaolong Wang, Zhe Zhao, Song Lai +5
While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinking remains understudied. This work evaluat…
ResearchEVO: An End-to-End Framework for Automated Scientific Discovery and Documentation
Zhe Zhao, Haibin Wen, Jiaming Ma +4
An important recurring pattern in scientific breakthroughs is a two-stage process: an initial phase of undirected experimentation that yields an unexpected finding, followed by a r…
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
Zhihui Yang, Yupei Wang, Kaijie Mo +2
Despite significant progress in multimodal language models (LMs), it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-on…
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals
Yupei Wang, Renfen Hu, Zhe Zhao
While current Automated Essay Scoring (AES) methods demonstrate high scoring agreement with human raters, their decision-making mechanisms are not fully understood. Our proposed me…