11 papers
Evaluation of OpenAI o1: Opportunities and Challenges of AGI
Tianyang Zhong, Zhengliang Liu, Yi Pan +73
This comprehensive study evaluates the performance of OpenAI's o1-preview large language model across a diverse array of complex reasoning tasks, spanning multiple domains, includi…
Build AI Assistants using Large Language Models and Agents to Enhance the Engineering Education of Biomechanics
Hanzhi Yan, Qin Lu, Xianqiao Wang +3
While large language models (LLMs) have demonstrated remarkable versatility across a wide range of general tasks, their effectiveness often diminishes in domain-specific applicatio…
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
Yun Wang, Zhaojun Ding, Xuansheng Wu +3
Automated scoring plays a crucial role in education by reducing the reliance on human raters, offering scalable and immediate evaluation of student work. While large language model…
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
Xuansheng Wu, Wenhao Yu, Xiaoming Zhai +1
Modern text classification methods heavily rely on contextual embeddings from large language models (LLMs). Compared to human-engineered features, these embeddings provide automati…
Understanding University Students' Use of Generative AI: The Roles of Demographics and Personality Traits
Newnew Deng, Edward Jiusi Liu, Xiaoming Zhai
The use of generative AI (GAI) among university students is rapidly increasing, yet empirical research on students' GAI use and the factors influencing it remains limited. To addre…
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
Xuansheng Wu, Jiayi Yuan, Wenlin Yao +2
Large language models (LLMs) excel at handling human queries, but they can occasionally generate flawed or unexpected responses. Understanding their internal states is crucial for…