papers

Publications (17)

cs.CY2026

Evaluating 21st-Century Competencies in Postsecondary Curricula with Large Language Models: Performance Benchmarking and Reasoning-Based Prompting Strategies

Zhen Xu, Xin Guan, Chenxi Shi +2

The growing emphasis on 21st-century competencies in postsecondary education, intensified by the transformative impact of generative AI, underscores the need to evaluate how these…

cs.CY2025

When the Past Misleads: Rethinking Training Data Expansion Under Temporal Distribution Shifts

Chengyuan Yao, Yunxuan Tang, Christopher Brooks +2

Predictive models are typically trained on historical data to predict future outcomes. While it is commonly assumed that training on more historical data would improve model perfor…

cs.CY2021

Should College Dropout Prediction Models Include Protected Attributes?

Renzhe Yu, Hansol Lee, René F. Kizilcec

Early identification of college dropouts can provide tremendous value for improving student success and institutional effectiveness, and predictive analytics are increasingly used…

cs.LG2022

A Robust Approach for the Decomposition of High-Energy-Consuming Industrial Loads with Deep Learning

Jia Cui, Yonghui Jin, Renzhe Yu +4

The knowledge of the users' electricity consumption pattern is an important coordinating mechanism between the utility company and the electricity consumers in terms of key decisio…

cs.CL2025

Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments

Li Siyan, Zhen Xu, Vethavikashini Chithrra Raghuram +3

Asynchronous learning environments (ALEs) are widely adopted for formal and informal learning, but timely and personalized support is often limited. In this context, Virtual Teachi…

cs.CY2025

From Course to Skill: Evaluating LLM Performance in Curricular Analytics

Zhen Xu, Xinjin Li, Yingqi Huan +2

Curricular analytics (CA) -- systematic analysis of curricula data to inform program and course refinement -- becomes an increasingly valuable tool to help institutions align acade…

cs.CY2024

Whose ChatGPT? Unveiling Real-World Educational Inequalities Introduced by Large Language Models

Renzhe Yu, Zhen Xu, Sky CH-Wang +1

The universal availability of ChatGPT and other similar tools since late 2022 has prompted tremendous public excitement and experimental effort about the potential of large languag…

econ.GN2026

AI-exposed jobs deteriorated before ChatGPT

Morgan R. Frank, Alireza Javadian Sabet, Lisa Simon +2

Public debate links worsening job prospects for AI-exposed occupations to the release of ChatGPT in late 2022. Using monthly U.S. unemployment insurance records, we measure occupat…

econ.GN2024

Course-Skill Atlas: A national longitudinal dataset of skills taught in U.S. higher education curricula

Alireza Javadian Sabet, Sarah H. Bana, Renzhe Yu +1

Higher education plays a critical role in driving an innovative economy by equipping students with knowledge and skills demanded by the workforce. While researchers and practitione…

cs.CY2025

Towards Fair and Privacy-Aware Transfer Learning for Educational Predictive Modeling: A Case Study on Retention Prediction in Community Colleges

Chengyuan Yao, Carmen Cortez, Renzhe Yu

Predictive analytics is widely used in learning analytics, but many resource-constrained institutions lack the capacity to develop their own models or rely on proprietary ones trai…

cs.CY2024

Temporal and Between-Group Variability in College Dropout Prediction

Dominik Glandorf, Hye Rin Lee, Gabe Avakian Orona +3

Large-scale administrative data is a common input in early warning systems for college dropout in higher education. Still, the terminology and methodology vary significantly across…

cs.CL2026

Enhancing LLM-Based Data Annotation with Error Decomposition

Zhen Xu, Vedant Khatri, Yijun Dai +4

Large language models offer a scalable alternative to human coding for data annotation tasks, enabling the scale-up of research across data-intensive domains. While LLMs are alread…

cs.CY2024

The Life Cycle of Large Language Models: A Review of Biases in Education

Jinsook Lee, Yann Hicke, Renzhe Yu +2

Large Language Models (LLMs) are increasingly adopted in educational contexts to provide personalized support to students and teachers. The unprecedented capacity of LLM-based appl…

cs.LG2024

Fairness Hub Technical Briefs: Definition and Detection of Distribution Shift

Nicolas Acevedo, Carmen Cortez, Chris Brooks +2

Distribution shift is a common situation in machine learning tasks, where the data used for training a model is different from the data the model is applied to in the real world. T…

cs.LG2023

Fairness Hub Technical Briefs: AUC Gap

Jinsook Lee, Chris Brooks, Renzhe Yu +1

To measure bias, we encourage teams to consider using AUC Gap: the absolute difference between the highest and lowest test AUC for subgroups (e.g., gender, race, SES, prior knowled…

cs.LG2023

Cross-Institutional Transfer Learning for Educational Models: Implications for Model Performance, Fairness, and Equity

Josh Gardner, Renzhe Yu, Quan Nguyen +2

Modern machine learning increasingly supports paradigms that are multi-institutional (using data from multiple institutions during training) or cross-institutional (using models fr…

cs.CY2021

Unsupervised Representations Predict Popularity of Peer-Shared Artifacts in an Online Learning Environment

Renzhe Yu, John Scott, Zachary A. Pardos

In online collaborative learning environments, students create content and construct their own knowledge through complex interactions over time. To facilitate effective social lear…